Volume 11 Advanced 5 sub-modules ~15 min read

Variation: Corners and OCV

Every delay so far has been one number. In silicon it is a spread: chips from the same wafer differ, hot chips are slower than cold ones, and two identical gates side by side never match exactly. This volume shows how timing analysis covers all of that - corners for the chip as a whole, derating for the variation within it - and how the tool gives back the doubt it would otherwise count twice.

You will learn
  • What PVT corners are, and why setup and hold are signed off at different ones
  • Why the worst RC corner depends on the length of the wire
  • How on-chip variation is covered by derating, and what it costs
  • How AOCV and POCV make derating less pessimistic
  • Where clock reconvergence pessimism comes from, and how CRPR removes it
You need

11.1 PVT corners

A PVT corner is one set of conditions - process, voltage, temperature - at which every delay on the chip is worked out. Setup is signed off at the slowest corner, hold at the fastest, because that is where each one fails.

Volume 02 showed a gate slowing down when the chip is hot, the supply is low, or the manufacturing came out slow. Those three move together across the whole chip, so a timing tool runs the whole analysis again at each extreme.

Corner Process Voltage Temperature Delays in this example
Slow (SS) Slow transistors Low Hot x1.45
Typical (TT) Typical Nominal Room x1.00
Fast (FF) Fast transistors High Cold x0.70

The same two paths at three corners

Take the adder path from Volume 04 and a short one-buffer path, on the 5 ns clock. Scale every delay by the corner's factor, and check both:

Corner Setup: adder path Setup: short path Hold: adder path Hold: short path
Slow 0.38 ns 4.64 ns 4.38 ns 0.12 ns
Typical 1.76 ns 4.70 ns 3.00 ns 0.06 ns
Fast 2.68 ns 4.74 ns 2.08 ns 0.02 ns

The library setup and hold times are kept fixed here to keep the sums clear; real libraries give them per corner too.

The long path's setup is worst at the slow corner, where its delay has grown most. The short path's hold is worst at the fast corner, where its data rushes through. Each check has its own worst corner.

Remember

Setup: sign off at the slow corner. Hold: sign off at the fast corner. A design is only finished when every check passes at every corner - usually many more than three.

When cold is slower than hot

On the smallest, lowest-voltage processes, the rule "hot is slow" can reverse. There, a cold transistor switches on more weakly, so the slowest corner may be cold. It is called temperature inversion, and it is one reason modern sign-off checks both temperature extremes at every voltage.

Common mistake

Signing off at the typical corner because it is what "most chips" will be. A chip that only works at typical fails in a hot car or on a sagging battery. Typical is for estimates, never for sign-off.

Quick check

Why is hold checked at the fast corner?

Show the answer

Answer: C. A hold violation is data arriving too soon. At the fast corner every delay is smallest, so data arrives earliest. The period plays no part in hold, so frequency has nothing to do with it.

11.2 RC corners

Wires vary too, and their resistance and capacitance do not always move together. The slowest RC corner for a short wire is not the slowest one for a long wire.

A metal line can come out a little thicker or thinner, wider or narrower, than drawn. A wider wire has less resistance but more capacitance to its neighbours. So RC corners come in pairs of opposites:

RC corner Resistance Capacitance 200 um wire 3000 um wire
Typical x1.00 x1.00 34.9 ps 1113.8 ps
Cworst x0.90 x1.15 39.1 ps 1199.0 ps
RCworst x1.20 x1.05 37.2 ps 1314.5 ps
Cbest x1.10 x0.85 30.6 ps 1007.9 ps
RCbest x0.80 x0.95 32.7 ps 926.7 ps

The wires are driven through 1 kOhm into 5 fF, as in Volume 02. On the short wire, the driver's charging of the wire's capacitance dominates, so the corner with the most capacitance is slowest: Cworst. On the long wire, the wire's own resistance matters most, so RCworst is slowest.

In plain words

There is no single "slow wire" corner. Short wires are slowest when capacitance is high; long wires are slowest when resistance is high. So sign-off checks every RC corner, combined with every PVT corner.

Quick check

A design's critical paths are mostly long wires between distant blocks. Which RC corner is most likely to be the worst for their setup?

Show the answer

Answer: A. Long wires are dominated by their own RC delay, which grows with the square of the length. RCworst raises the resistance most, so it is usually the worst corner for them.

11.3 On-chip variation and derating

Corners treat the whole chip as uniformly slow or fast. But gates on one chip also differ from each other. On-chip variation is covered by derating: pushing some delays slower and others faster, to assume the worst mix.

Even at the slow corner, the gates on one chip are not all equally slow. One corner of the die may be a little hotter; one cell may have come out a little different. Timing has to assume the unlucky mix.

For setup, the unlucky mix is: the launch clock and the data path slow, the capture clock fast. For hold it is the reverse. A derate of 5% means multiplying the "slow" side by 1.05 and the "fast" side by 0.95.

The adder path at the slow corner, derated

Derate Launch clock + clock-to-Q + data Capture clock Setup slack
None 5.73 ns 1.28 ns 0.38 ns
+/-5% 6.01 ns 1.21 ns 0.03 ns
+/-8% 6.19 ns 1.17 ns -0.18 ns

The path that passed at the slow corner fails once on-chip variation is added. That is the whole point: the corner covers the chip being slow, the derate covers it being unevenly slow.


set_timing_derate -late  1.05
set_timing_derate -early 0.95
Common mistake

Applying the derate to the whole slack instead of to the delays. A 5% derate is not "5% less slack". It stretches every delay on one side and shrinks every delay on the other, so long paths and deep clock trees pay much more than short ones.

Quick check

A data path of 2.00 ns is given a late derate of 8%. What delay does the setup check use?

Show the answer

Answer: D. A late derate of 8% multiplies the delay by 1.08: 2.00 x 1.08 = 2.16 ns. The early derate, 0.92, would be used for the same path's hold check.

11.4 AOCV and POCV, simply

A flat derate assumes every gate on a path is unlucky at once. Real variation is partly random, and random differences tend to cancel over many gates. AOCV and POCV use that to remove pessimism.

AOCV: deeper paths get a smaller derate

If each gate's variation is random, a path of many gates averages it out. So AOCV looks up the derate from a table by the path's logic depth. A typical shape shrinks with the square root of the depth:

Logic depth Late derate
1 +15.0%
2 +10.6%
4 +7.5%
8 +5.3%
16 +3.8%
32 +2.7%
An AOCV derate shrinking as logic depth grows 0 4 8 12 16 20 24 28 32 0 4 8 12 16 logic depth (gates on the path) late derate (%)
Figure 11.1 - A single gate might be 15% slow. Over 32 gates the random parts mostly cancel, so the path as a whole needs only a 2.7% derate. This example uses 1 + 0.15 divided by the square root of the depth.

POCV: add the spreads, not the worst cases

POCV goes one step further. Each cell has an average delay and a spread, its sigma. The tool adds the averages, and combines the spreads statistically - as the square root of the sum of their squares.

Take the adder path's five cells, each with a sigma of 5% of its delay:

Method Path delay
Average 2.890 ns
Every cell at +3 sigma at once (flat) 3.323 ns
The whole path at +3 sigma (POCV) 3.211 ns

POCV removes 0.113 ns of pessimism on this one path. It removes more on paths with many similar cells: eight cells of 0.30 ns come out at 2.760 ns flat, but 2.527 ns with POCV.

In plain words

Flat derating assumes every cell is unlucky at once. POCV assumes they are unlucky independently, which is what really happens. The result is still safe, but less wasteful.

Quick check

With the AOCV table shape above, what late derate does a path nine gates deep get?

Show the answer

Answer: B. The derate is 1 + 0.15 divided by the square root of the depth. The square root of 9 is 3, and 0.15 / 3 = 0.05, so the late derate is +5.0%.

11.5 Clock reconvergence pessimism (CRPR)

Suppose the launch and capture clocks share part of the clock tree. Derating then treats that shared part as slow on one side and fast on the other, at the same moment. That is impossible, and CRPR gives the difference back.

A clock tree with a shared trunk splitting into a branch to the launch flop and a branch to the capture flop clock in shared: 0.60 ns 0.25 ns 0.28 ns launch flop capture flop
Figure 11.2 - The first 0.60 ns of the clock tree is shared by both flip-flops. Any variation there delays both clocks alike, so it cannot make one late and the other early. Only the two branches after the split can really differ.

Where the pessimism comes from

Derate the clock tree by 5% each way, for a setup check:

  1. Launch clock, made late: (0.60 + 0.25) x 1.05 = 0.8925 ns.
  2. Capture clock, made early: (0.60 + 0.28) x 0.95 = 0.8360 ns.
  3. The shared 0.60 ns was counted as 0.630 ns on one side and 0.570 ns on the other. One piece of wire cannot be both at once.
  4. CRPR credit: 0.630 - 0.570 = 0.060 ns, given back to the check.
Adder path at 5 ns, derated 5% Without CRPR With CRPR
Setup slack 1.52 ns 1.58 ns
Hold slack 1.07 ns 1.13 ns

In the report


Startpoint: u_launch
Endpoint:   u_capture
Path Group: clk
Path Type:  max

  Point                                              Incr     Path
  ----------------------------------------------------------------
  clock clk (rise edge)                              0.00     0.00
  clock network delay (propagated)                   0.89     0.89
  u_launch/CK                                        0.00     0.89
  u_launch/Q (clock-to-Q)                            0.22     1.11
  logic                                              3.03     4.15
  u_capture/D                                        0.00     4.15
  data arrival time                                           4.15

  clock clk (rise edge)                              5.00     5.00
  clock network delay (propagated)                   0.84     5.84
  u_capture/CK                                       0.00     5.84
  clock reconvergence pessimism                      0.06     5.90
  clock uncertainty                                 -0.08     5.82
  library setup time                                -0.09     5.73
  data required time                                          5.73
  ----------------------------------------------------------------
  data required time                                          5.73
  data arrival time                                          -4.15
  ----------------------------------------------------------------
  slack (MET)                                                 1.58

The clock reconvergence pessimism line sits on the required side, and adds the 0.06 ns back. The clock and logic lines above it are already derated: 3.03 ns of logic is 2.89 ns x 1.05.

Remember

The more of the clock tree two flip-flops share, the more CRPR gives back. Flip-flops that talk to each other are best placed close together on the same clock branch - it helps skew and CRPR alike.

Quick check

Two flip-flops share 1.00 ns of clock tree, derated by 10% each way. How much does CRPR give back?

Show the answer

Answer: A. The shared part is counted as 1.10 ns on the late side and 0.90 ns on the early side. The difference, 0.20 ns, is impossible pessimism, and CRPR removes it.

What you learned

Key words from this volume

Every word below has a plain-English entry in the glossary.

Practice

Practice 1

A faster clock at the slow corner

At the slow corner the adder path has 0.38 ns of setup slack on a 5 ns clock. Is it safe to run it at 4.2 ns (238 MHz)?

Show the solution

No. Taking 0.80 ns off the period takes 0.80 ns off the slack: 0.38 - 0.80 = -0.42 ns. The path would fail at the slow corner, before any on-chip variation is even added.

Practice 2

Which corner for which path?

A design has three worrying paths: a long path between two far-apart blocks, a short path with one gate, and a medium path of many small gates. Which corner is each most likely to fail at?

Show the solution

The long, wire-dominated path is most likely to fail setup at the slow PVT corner with RCworst wires. The short path is most likely to fail hold at the fast PVT corner. The medium path of many gates is gate-dominated. It is most likely to fail setup at the slow PVT corner with Cworst wires, where each gate drives the most capacitance.

In practice the tool checks them all everywhere. But knowing where each one is weakest tells you which report to open first.

Interview corner

Interview question 1

Explain CRPR

"What is clock reconvergence pessimism, and why is it removed?"

Show the solution

"With on-chip variation derating, the launch clock path is made late and the capture clock path early for setup. If the two share a common segment of the clock tree, that segment is derated late on one side and early on the other in the same check. But one physical buffer cannot be both. The pessimism equals the common segment's delay times the difference between the late and early derates. CRPR subtracts it from the requirement, which can recover tens of picoseconds on deep clock trees."

Interview question 2

Corners versus OCV

"If you already check the slow corner, why do you need OCV as well?"

Show the solution

"A corner assumes the whole die is uniformly slow or fast. OCV covers variation within one die at that corner: local process differences, voltage drop in one region, a hot spot. Those can make the launch path slower than the capture clock even at the same corner. So the corner covers the global case, and derating - flat, AOCV or POCV - covers the local one on top."

Volume 12 turns to a different kind of variation: a delay that changes because a neighbouring wire switches at the same moment.