STA Problem Vault and Revision
This is the revision volume. It starts with the whole course on one sheet, then gives you forty numerical problems to solve, twenty questions to answer out loud, five timing reports to read like an engineer, and a deck of flashcards. Every answer was worked out by the same timing model that produced the numbers in the other fifteen volumes.
- Every formula in the course, on one sheet, with the volume that explains it
- How to solve setup, hold, clocking, exception, latch, I/O and variation problems quickly
- Model answers to the twenty concept questions interviewers ask most
- How to read a timing report and find what is wrong in under a minute
- The key facts of the course, as flashcards
- Volumes 00 to 14, or at least the ones whose problems you want to try
- A calculator and some paper
15.1 Formula sheet
The whole course fits on two pages: the sums, and the one idea behind each. Read this sheet the evening before an exam or an interview.
If a line feels unfamiliar, the first column links back to the volume that explains it.
The ideas, volume by volume
| Volume | Remember this |
|---|---|
| 00 Start here | A clock edge makes every flip-flop copy its input. Between edges, data must cross the logic in time. T (ns) = 1000 / f (MHz). |
| 01 Why timing analysis | Setup: data must arrive before the next edge. Hold: it must not arrive so soon that it spoils the current one. STA checks every path without test patterns. |
| 02 Where delay comes from | A gate's delay is its own delay plus 0.69 R C. Wires add a term that grows with the square of their length. Libraries store delay in tables of slew and load. |
| 03 Timing paths | A path runs from a start point to an end point: input or clock pin, to a data pin or output. There are four types: in-to-reg, reg-to-reg, reg-to-out and in-to-out. |
| 04 Setup | Slack = required - arrival. The worst setup path sets the fastest clock. |
| 05 Hold | Hold uses the fastest delays, and the clock period is not in it. A slower clock cannot fix hold. |
| 06 The clock | Latency, skew, jitter and uncertainty. Skew that helps setup hurts hold on the same path. |
| 07 Every clocking case | Half-cycle paths get part of a period. Two related clocks get the GCD of their periods. Unrelated clocks must not be timed as if related. |
| 08 Exceptions | False paths are never checked. A multicycle setup needs a matching hold. Max delay bounds a crossing. |
| 09 Latches | A latch is open for part of the cycle. Data that arrives while it is open borrows time from the next stage. |
| 10 Inputs and outputs | Input and output delays describe the chip next door. DDR moves a bit on each clock edge. |
| 11 Corners and OCV | Setup is usually worst at the slow corner, hold at the fast one. Derating models variation on one chip; CRPR removes the part it double-counts. |
| 12 Signal integrity | A switching neighbour slows or speeds a net through coupling. It can also cause a glitch. |
| 13 SDC | Clocks, generated clocks, I/O delays, exceptions and design rules, written in SDC. An unconstrained path is not checked at all. |
| 14 Closure | Fix setup with the biggest lever first. Fix hold after the clock tree exists. WNS and TNS track progress. |
The two sums that carry the course
setup slack = required - arrival
arrival = launch edge + launch clock + clock-to-Q (slowest) + logic (slowest)
required = capture edge + capture clock - uncertainty - setup time (+ CRPR)
on one line: T + skew - clock-to-Q - logic - setup time - uncertainty
hold slack = arrival - required
on one line: clock-to-Q (fastest) + logic (fastest) - hold time - skew - hold uncertainty
skew = capture clock arrival - launch clock arrival
f_max = 1000 / (T - setup slack) in MHz, when T is in ns
A positive skew helps setup and hurts hold. The period T appears in the setup sum and not in the hold sum.
Numbers and formulas
| What | Formula | Volume |
|---|---|---|
| Period and frequency | T (ns) = 1000 / f (MHz) | 00 |
| Gate delay | own delay + 0.69 x R x C, where kOhm x fF = ps | 02 |
| Slew, 10% to 90% | 2.2 x R x C | 02 |
| Wire delay | 0.69 R_d (C_w + C_L) + 0.38 R_w C_w + 0.69 R_w C_L | 02 |
| Table lookup | interpolate along the load, then along the slew | 02 |
| Setup uncertainty | before the clock tree: jitter + skew estimate + margin; after it: jitter + margin | 06 |
| Half-cycle path | rise to fall gets duty x T; fall to rise gets the rest of the period | 07 |
| Two related clocks | the tightest setup check is the GCD of their periods | 07 |
| Multicycle | -setup N moves the capture edge N - 1 periods later; -hold N - 1 brings hold back | 08 |
| Latch | most it can lend = time open - setup; borrowed = arrival - opening edge | 09 |
| Input delay | -max = their clock-to-out + trace, both slowest; -min = both fastest | 10 |
| Output delay | -max = trace (slowest) + their setup; -min = trace (fastest) - their hold | 10 |
| DDR, strobe centred | bit time = T / 2; setup margin = half a bit - skew - setup | 10 |
| OCV | late delays x (1 + d); early delays x (1 - d) | 11 |
| CRPR | the shared clock, late minus early | 11 |
| AOCV, the course's table | late derate 1 + 0.15 / sqrt(depth) | 11 |
| POCV | mean + 3 x sqrt(sigma1² + sigma2² + ...) | 11 |
| Crosstalk delay | 0.69 R (C_g + k C_c): k = 0 same way, 1 quiet, 2 opposite | 12 |
| Glitch, upper bound | Vdd x C_c / (C_c + C_g) | 12 |
| Hold fix | cells = violation / one cell's fastest delay, rounded up | 14 |
| WNS and TNS | the worst slack; the sum of the failing slacks | 14 |
0.69 is ln 2 = 0.693, and 2.2 is ln 9 = 2.197. The answers in this volume use the exact values.
A path has setup slack -0.15 ns at 400 MHz. What is its maximum frequency?
Show the answer
Answer: B. 400 MHz is 2.50 ns. The path needs 0.15 ns more, so its shortest period is 2.50 + 0.15 = 2.65 ns, and 1000 / 2.65 = 377.4 MHz. 425.5 MHz comes from subtracting the 0.15 instead of adding it.
15.2 40 numerical problems, solved
Forty problems, in course order. Try each one on paper before you open the solution - the effort of getting stuck is where the learning happens.
All of them were written for this course, in the style of university exams, GATE and chip-design interviews. Times are in ns unless a problem says otherwise, and there is no clock tree unless one is given.
Units and delay
1. A period in two units
A chip runs at 750 MHz. What is its clock period in ns and in ps? If the flip-flops use 0.15 ns of each cycle, how many gates of 0.09 ns fit in the rest?
Show the solution
Period = 1000 / 750 = 1.333 ns = 1333 ps.
The logic may use 1.333 - 0.15 = 1.183 ns. That is 1.183 / 0.09 = 13.15 gate delays, so 13 gates fit. A fourteenth would not.
2. One gate, one load
A gate has an own delay of 12 ps and a drive resistance of 1.5 kOhm. It drives 8 fF. What are its delay and its output slew?
Show the solution
R x C = 1.5 kOhm x 8 fF = 12 ps, since kOhm x fF gives ps.
Delay = 12 + 0.693 x 12 = 12 + 8.3 = 20.3 ps. Slew = 2.197 x 12 = 26.4 ps.
3. A long wire
A 1 kOhm driver sends a signal down 2000 µm of wire to a 5 fF load. The wire has 1 Ohm and 0.2 fF per µm. What is the delay? What happens to the wire's own term at 4000 µm?
Show the solution
The wire has R_w = 2.0 kOhm and C_w = 400 fF.
| Term | Sum | Delay |
|---|---|---|
| Driver | 0.693 x 1 x (400 + 5) | 281 ps |
| Wire | 0.38 x 2.0 x 400 | 304 ps |
| Load | 0.693 x 2.0 x 5 | 7 ps |
| Total | 592 ps |
At 4000 µm the wire term is 0.38 x 4.0 x 800 = 1216 ps, four times as much. Doubling the length doubles both R_w and C_w, so their product grows by four.
4. Where to put repeaters
A 3000 µm wire of the same kind is cut into equal pieces, each driven by a buffer (15 ps, 1.0 kOhm, 5 fF input). The last piece drives 5 fF. How many pieces give the least delay?
Show the solution
The timing model tries every count:
| Pieces | Each piece | Total delay |
|---|---|---|
| 1 | 3000 µm | 1128.8 ps |
| 2 | 1500 µm | 805.2 ps |
| 3 | 1000 µm | 709.7 ps |
| 4 | 750 µm | 671.1 ps |
| 6 | 500 µm | 651.1 ps |
| 8 | 375 µm | 659.5 ps |
| 12 | 250 µm | 704.9 ps |
Six pieces are best. Cutting the wire shrinks the square-law term, but each extra buffer adds its own delay. Past six, the buffers cost more than they save.
5. A table lookup
Find the NAND2_X1 delay for an input slew of 100 ps and a load of 8 fF. The four table entries around that point are:
| Slew \ Load | 4 fF | 16 fF |
|---|---|---|
| 50 ps | 29.1 ps | 54.4 ps |
| 150 ps | 54.3 ps | 80.3 ps |
Show the solution
8 fF is 4/12 = 1/3 of the way from 4 to 16 fF. 100 ps is halfway from 50 to 150 ps.
- Along the load at 50 ps: 29.1 + (54.4 - 29.1) / 3 = 37.53 ps.
- Along the load at 150 ps: 54.3 + (80.3 - 54.3) / 3 = 62.97 ps.
- Along the slew: 37.53 + (62.97 - 37.53) x 0.5 = 50.25 ps.
Setup
6. A plain setup check
At 250 MHz, a path has clock-to-Q 0.25, 3.20 of logic, setup time 0.10 and uncertainty 0.10. What is its setup slack?
Show the solution
Arrival = 0.25 + 3.20 = 3.45. Required = 4.00 - 0.10 - 0.10 = 3.80.
Slack = 3.80 - 3.45 = 0.35 ns. It passes.
7. A helpful clock tree
A 3 ns clock reaches the launch flip-flop at 0.90 and the capture flip-flop at 1.05. Clock-to-Q is 0.18, logic 2.70, setup 0.08 and uncertainty 0.05. What is the setup slack?
Show the solution
Arrival = 0.90 + 0.18 + 2.70 = 3.78. Required = 3.00 + 1.05 - 0.05 - 0.08 = 3.92.
Slack = 0.14 ns. The skew is 1.05 - 0.90 = +0.15, and all of it helps setup.
8. Maximum frequency
A path has clock-to-Q 0.20, logic 4.35, setup 0.12 and uncertainty 0.08. What is its slack at 5 ns, and how fast can it run?
Show the solution
Slack = 5.00 - 0.20 - 4.35 - 0.12 - 0.08 = 0.25 ns.
Shortest period = 5.00 - 0.25 = 4.75 ns, so f_max = 1000 / 4.75 = 210.5 MHz.
9. A logic budget
A design must run at 800 MHz. Its flip-flops have clock-to-Q 0.09 and setup 0.05, and the uncertainty is 0.06. How many levels of 0.08 ns logic fit between two flip-flops?
Show the solution
800 MHz is 1.25 ns. The budget is 1.25 - 0.09 - 0.05 - 0.06 = 1.05 ns.
1.05 / 0.08 = 13.1, so 13 levels fit, using 1.04 ns with 0.01 ns to spare. Fourteen levels would fail by 0.07 ns.
10. An unhelpful clock tree
A 5 ns clock reaches the launch flip-flop at 1.10 and the capture flip-flop at 0.95. Clock-to-Q is 0.22, logic 4.50 and setup 0.10. What is the setup slack?
Show the solution
The skew is 0.95 - 1.10 = -0.15. Slack = 5.00 - 0.15 - 0.22 - 4.50 - 0.10 = 0.03 ns.
It still passes, but only just. With a balanced tree it would have had 0.18 ns.
Hold
11. A plain hold check
A path has fastest clock-to-Q 0.12, fastest logic 0.05 and hold time 0.08. There is no skew. What is its hold slack?
Show the solution
Arrival = 0.12 + 0.05 = 0.17. Required = 0.08.
Hold slack = 0.17 - 0.08 = 0.09 ns. It passes.
12. A late capture clock
The clock reaches the launch flip-flop at 0.60 and the capture flip-flop at 0.85. The fastest clock-to-Q is 0.10, the fastest logic 0.08, the hold time 0.05 and the hold uncertainty 0.03. What is the hold slack?
Show the solution
Arrival = 0.60 + 0.10 + 0.08 = 0.78. Required = 0.85 + 0.03 + 0.05 = 0.93.
Hold slack = 0.78 - 0.93 = -0.15 ns. It fails, because the capture clock arrives 0.25 ns late.
13. Fixing it with delay cells
Fix problem 12 with delay cells of 0.04 ns at the fast corner and 0.07 ns at the slow corner. The path has 0.79 ns of setup slack. How many cells, and what is left of each slack?
Show the solution
0.15 / 0.04 = 3.75, so 4 cells. They add 4 x 0.04 = 0.16 ns, so hold slack becomes 0.01 ns.
At the slow corner they add 4 x 0.07 = 0.28 ns, so setup slack falls to 0.79 - 0.28 = 0.51 ns. Both pass.
14. The skew window
On a 2 ns clock, a path has clock-to-Q 0.08 fastest and 0.12 slowest, and logic 0.10 fastest and 1.60 slowest. Setup is 0.06 and hold 0.05. What range of skew lets both checks pass?
Show the solution
Setup: 2.00 + skew - 0.12 - 1.60 - 0.06 must be at least 0, so skew must be at least -0.22 ns.
Hold: 0.08 + 0.10 - 0.05 - skew must be at least 0, so skew must be at most +0.13 ns.
The window is 0.35 ns wide. At either end, one slack is 0.00 and the other is 0.35.
The clock
15. Duty cycle matters
A path launches on the rising edge of a 6 ns clock and is captured on the falling edge. Clock-to-Q is 0.15, logic 2.10, setup 0.08. What is the slack at 50%, 40% and 35% duty?
Show the solution
Rise to fall gets duty x T. The path needs 0.15 + 2.10 + 0.08 = 2.33 ns.
| Duty | Time allowed | Slack |
|---|---|---|
| 50% | 3.00 ns | 0.67 ns |
| 40% | 2.40 ns | 0.07 ns |
| 35% | 2.10 ns | -0.23 ns |
16. Building the uncertainty
A clock has 0.06 ns of jitter. Before the clock tree is built, the team allows 0.15 ns for skew and 0.04 ns of margin. What setup and hold uncertainty should be used before and after the tree?
Show the solution
| Setup | Hold | |
|---|---|---|
| Before the tree | 0.06 + 0.15 + 0.04 = 0.25 | 0.15 + 0.04 = 0.19 |
| After the tree | 0.06 + 0.04 = 0.10 | 0.04 |
Hold leaves out jitter, because it compares data against the same edge. After the tree, the real skew is in the report, so the estimate goes.
17. Global and local skew
The clock reaches four flip-flops at FF1 1.02, FF2 1.10, FF3 0.96 and FF4 1.15. Data flows FF1 to FF2 to FF3 to FF4. Find the global skew and each local skew.
Show the solution
Global skew = latest - earliest = 1.15 - 0.96 = 0.19 ns.
Local skew is capture minus launch: FF1 to FF2 +0.08, FF2 to FF3 -0.14, FF3 to FF4 +0.19. The FF2 to FF3 path loses 0.14 ns of setup time.
18. Useful skew
On a 4 ns clock, stage A runs FF1 to FF2 and stage B runs FF2 to FF3. Clock-to-Q is 0.12 to 0.20, setup 0.08, hold 0.04. A's logic is 0.40 to 3.92; B's is 0.30 to 3.22. What happens if FF2's clock is delayed by 0.25 ns?
Show the solution
| A setup | A hold | B setup | B hold | |
|---|---|---|---|---|
| Balanced clock | -0.20 | 0.48 | 0.50 | 0.38 |
| FF2 clock 0.25 later | 0.05 | 0.23 | 0.25 | 0.63 |
A gains 0.25 ns of setup and B loses it. Everything now passes. A's hold slack also dropped by 0.25, so it must still be checked.
Clocking cases
19. A half-cycle path
On a 10 ns clock with 50% duty, data launches on the rising edge and is captured on the falling edge. Clock-to-Q is 0.18 to 0.25, logic 0.60 to 4.20, setup 0.10, hold 0.05. Find both slacks.
Show the solution
Setup: launch at 0, capture at 5. Slack = 5.00 - 0.25 - 4.20 - 0.10 = 0.45 ns.
Hold: the data must not spoil the capture made at the falling edge before, at -5. Slack = 0.18 + 0.60 - (-5.00 + 0.05) = 5.73 ns. Half-cycle paths are hard on setup and easy on hold.
20. 100 MHz to 125 MHz
Data goes from a 10 ns clock to an 8 ns clock. Both rise together at 0. Clock-to-Q is 0.20, logic 1.50, setup 0.10. What is the tightest setup check, and its slack?
Show the solution
The launch edges are 0, 10, 20 and 30; after each, the next capture edge is 8, 16, 24 and 32. The gaps are 8, 6, 4 and 2 ns, then the pattern repeats every 40 ns.
2 ns is the GCD of 10 and 8. Slack = 2.00 - 0.20 - 1.50 - 0.10 = 0.20 ns.
21. A generated clock from edges
A clock is generated from an 8 ns master with -edges {1 5 7}. What are its period, fall time and duty
cycle? Compare -divide_by 2.
Show the solution
The master's edges are numbered from 1: edge 1 rises at 0, edge 2 falls at 4, edge 3 rises at 8, and so on. Edge 5 is at 16 and edge 7 at 24.
So the clock rises at 0, falls at 16 and rises again at 24: period 24 ns, 67% duty. -divide_by 2
is -edges {1 3 5}: period 16 ns, falling at 8, 50% duty.
22. A clock-gating check
An AND gate gates an 8 ns clock that is high from 0 to 4. The enable comes from a flip-flop through 0.50 ns of logic, with clock-to-Q 0.20. Compare an enable flip-flop clocked on the rising edge with one on the falling edge.
Show the solution
The enable may only change while the clock is low, from 4 to 8.
- Rising-edge flip-flop: the enable changes at 0.70, while the clock is high. Gating hold slack is 0.70 - 4.00 = -3.30 ns - the gated clock would glitch.
- Falling-edge flip-flop: it changes at 4.00 + 0.70 = 4.70. Gating setup slack is 8.00 - 4.70 = 3.30 ns, and gating hold slack 4.70 - 4.00 = 0.70 ns. Both pass.
Exceptions
23. A multicycle path, half done
On a 10 ns clock, a path takes 17.50 ns (1.00 fastest). Clock-to-Q is 0.20 to 0.30, setup 0.10, hold
0.05. Find both slacks with -setup 2 alone, then with -hold 1 added.
Show the solution
-setup 2 moves capture to 20: setup slack = 20.00 - 0.30 - 17.50 - 0.10 = 2.10 ns.
But hold is dragged along to the edge at 10. Hold slack = 0.20 + 1.00 - 10.00 - 0.05 = -8.85 ns.
Adding -hold 1 moves it back to 0, giving 1.15 ns.
24. Three cycles on a fast clock
On a 5 ns clock, a path takes 13.80 ns (0.90 fastest), with clock-to-Q 0.20 to 0.30, setup 0.10 and
hold 0.05. It has -setup 3 -hold 2. Find both slacks.
Show the solution
Setup is checked at 15: slack = 15.00 - 0.30 - 13.80 - 0.10 = 0.80 ns.
Hold is back at 0: slack = 0.20 + 0.90 - 0.05 = 1.05 ns.
25. Fast to slow, counted at the start
Data goes from a 5 ns clock to a 15 ns clock. The logic takes 12.00 ns (0.80 fastest), clock-to-Q 0.20,
setup 0.10, hold 0.05. Find the slacks with no exception, with -setup 3 -start, and with -hold 2 -start added.
Show the solution
| Constraint | Setup edges | Setup slack | Hold edges | Hold slack |
|---|---|---|---|---|
| None | 10 to 15 | -7.30 | 0 to 0 | 0.95 |
-setup 3 -start |
0 to 15 | 2.70 | 5 to 15 | -9.05 |
-setup 3 -hold 2 -start |
0 to 15 | 2.70 | 15 to 15 | 0.95 |
-start counts in launch clock periods, the fast clock here. Moving the launch edge two periods earlier
gives the full 15 ns.
26. A max delay on a crossing
A crossing between unrelated clocks has set_max_delay -datapath_only 2.50. Clock-to-Q is 0.25, the
route 2.05 and setup 0.10. The clock trees are 0.70 and 0.40. What is the slack?
Show the solution
-datapath_only ignores both clock trees. Slack = 2.50 - 0.25 - 2.05 - 0.10 = 0.10 ns.
Latches
27. Borrowing time
A latch on an 8 ns clock is open from 8 to 12, with setup 0.10. Data arrives at 10.30. How much does it borrow, what is its slack, and what is the most it could borrow?
Show the solution
It arrives 10.30 - 8.00 = 2.30 ns after the latch opened. Slack = 12.00 - 0.10 - 10.30 = 1.60 ns.
The most it can lend is the open time less setup: 4.00 - 0.10 = 3.90 ns.
28. A narrow window
A 10 ns clock has 40% duty, and a latch is open from 10 to 14 with setup 0.12. Data arrives at 14.00. Does it make it?
Show the solution
The latch needs the data by 14.00 - 0.12 = 13.88. Slack = -0.12 ns, so it fails.
The most this latch can lend is 4.00 - 0.12 = 3.88 ns. At 50% duty it would have been open until 15, and the same data would have passed.
29. A latch pipeline
On an 8 ns clock, a flip-flop launches at 0 (clock-to-Q 0.25). Latch 1 is open from 4 to 8, latch 2 from 8 to 12, and a flip-flop captures at 16. The logic is 5.00, 3.60 and 6.50 ns. Latches pass data through in 0.15 ns; every setup is 0.10. Walk the data through.
Show the solution
| Stage | Logic | Arrives | Borrowed | Slack | Leaves |
|---|---|---|---|---|---|
| 1 | 5.00 | 5.25 | 1.25 | 2.65 | 5.40 |
| 2 | 3.60 | 9.00 | 1.00 | 2.90 | 9.15 |
| 3 | 6.50 | 15.65 | - | 0.25 | - |
Every stage passes. With a flip-flop at 4 ns in place of latch 1, stage 1 would fail by 1.35 ns. The latches let 15.10 ns of uneven logic share 16 ns.
Inputs and outputs
30. An input from a datasheet
A sending chip has clock-to-out 1.2 to 3.5 ns, and the trace takes 0.3 to 0.6 ns. Inside, the input goes through 1.30 to 4.80 ns of logic to a flip-flop whose clock tree is 0.50, with setup 0.10 and hold 0.05. The clock is 10 ns. Find the input delays and both slacks.
Show the solution
Input delay -max = 3.5 + 0.6 = 4.10; -min = 1.2 + 0.3 = 1.50.
Setup: arrival 4.10 + 4.80 = 8.90; required 10.00 + 0.50 - 0.10 = 10.40; slack 1.50 ns.
Hold: arrival 1.50 + 1.30 = 2.80; required 0.50 + 0.05 = 0.55; slack 2.25 ns.
31. An output to a datasheet
A receiving chip needs setup 1.8 and hold 0.6, and the trace takes 0.5 to 0.8. Inside, the clock tree is 0.50, clock-to-Q 0.18 to 0.25, and logic plus pad 1.10 to 3.00. The clock is 8 ns. Find the output delays and both slacks.
Show the solution
Output delay -max = 0.8 + 1.8 = 2.60; -min = 0.5 - 0.6 = -0.10.
Setup: arrival 0.50 + 0.25 + 3.00 = 3.75; required 8.00 - 2.60 = 5.40; slack 1.65 ns.
Hold: arrival 0.50 + 0.18 + 1.10 = 1.78; required 0 - (-0.10) = 0.10; slack 1.68 ns.
32. A board-level limit
Two chips share one board clock. The sender's clock-to-out is 3.0, the trace 0.9, the board clock skew 0.4, and the receiver needs 1.1 for its input path and setup. What is the fastest clock?
Show the solution
The period must cover 3.0 + 0.9 + 0.4 + 1.1 = 5.4 ns, so the fastest clock is 1000 / 5.4 = 185.2 MHz. Every part of the sum is fixed by the board, which is why fast interfaces forward their own clock.
33. A DDR margin
A DDR bus runs at 300 MHz with the strobe centred in each bit. Data-to-strobe skew is up to 0.25 ns and the receiver's setup is 0.15 ns. What is the setup margin?
Show the solution
Period = 3.333 ns, so each bit lasts 1.667 ns. The strobe sits half a bit, 0.833 ns, after the data changes.
Margin = 0.833 - 0.25 - 0.15 = 0.433 ns.
Variation
34. Derating a path
On a 5 ns clock, both flip-flops' clock trees are 1.00. Clock-to-Q is 0.20, logic 3.00, setup 0.10 and uncertainty 0.05. Find the setup slack without derating, then with ±6% OCV.
Show the solution
Without derating: slack = 5.00 - 0.20 - 3.00 - 0.10 - 0.05 = 1.65 ns.
With ±6%, the launch side is late: arrival = (1.00 + 0.20 + 3.00) x 1.06 = 4.452. The capture clock is early: required = 5.00 + 1.00 x 0.94 - 0.05 - 0.10 = 5.790. Slack = 1.34 ns.
35. Giving back the pessimism
In problem 34, the first 0.80 ns of both clock trees is shared. How much CRPR is due, and what is the slack?
Show the solution
The shared 0.80 was counted late (0.848) on the launch side and early (0.752) on the capture side. One wire cannot be both.
CRPR = 0.848 - 0.752 = 0.096 ns, so the slack becomes 1.338 + 0.096 = 1.43 ns.
36. Depth-based derating
Using the course's AOCV table, late derate = 1 + 0.15 / sqrt(depth), what does a 2.00 ns path become at depth 1, 4 and 25?
Show the solution
| Depth | Derate | 2.00 ns becomes |
|---|---|---|
| 1 | +15.0% | 2.300 ns |
| 4 | +7.5% | 2.150 ns |
| 25 | +3.0% | 2.060 ns |
A deep path gets a gentler derate. Its cells' random variations partly cancel each other.
37. Statistical timing
A path has four cells of 0.50 ns, each with a sigma of 5% of its delay. Compare every cell at +3 sigma with the path at +3 sigma (POCV).
Show the solution
Each cell's sigma is 0.025 ns.
- Every cell at +3 sigma: 4 x (0.50 + 0.075) = 2.300 ns.
- POCV: the path's sigma is sqrt(4) x 0.025 = 0.050, so 2.00 + 3 x 0.050 = 2.150 ns.
POCV removes 0.150 ns of pessimism. All four cells are very unlikely to be slow at once.
Crosstalk and closure
38. A victim net
A net is driven through 1.2 kOhm. It has 25 fF to ground and 10 fF to a neighbour. Find its delay with the neighbour quiet, switching the opposite way and switching the same way.
Show the solution
| Neighbour | Capacitance counted | Delay |
|---|---|---|
| Quiet | 25 + 10 = 35 fF | 29.1 ps |
| Opposite way | 25 + 2 x 10 = 45 fF | 37.4 ps (+8.3) |
| Same way | 25 + 0 = 25 fF | 20.8 ps (-8.3) |
Opposite switching hurts setup; same-way switching hurts hold.
39. A glitch
A quiet net has 24 fF to ground and 6 fF to a neighbour that switches on a 1.0 V supply. What is the upper bound on the glitch?
Show the solution
1.0 x 6 / (6 + 24) = 0.200 V, or 20% of the supply. A real driver holds the net and fights the bump, so the true peak is lower.
40. WNS and TNS
Five paths have setup slacks of -0.22, -0.08, +0.05, -0.31 and +0.40 ns. Find the WNS, the TNS and the number of failing paths. What are they after the two worst are fixed?
Show the solution
WNS = -0.31 ns. TNS = -0.22 - 0.08 - 0.31 = -0.61 ns, with 3 paths failing.
After fixing the two worst, only the -0.08 path fails: WNS and TNS are both -0.08 ns.
A path has hold slack -0.10 ns at 100 MHz. What is its hold slack at 200 MHz?
Show the answer
Answer: A. The clock period is not in the hold sum, so changing the frequency changes nothing. The same violation fails at every speed, which is why a hold failure on silicon cannot be fixed by slowing the clock.
15.3 20 concept questions
Interviewers ask about timing to see whether you understand what the hardware does, not whether you can recite a formula. A short answer with a reason beats a long one.
Read each question, answer it out loud, and only then compare your answer with the model. The model answers are short on purpose.
What is STA?
"What is static timing analysis, and why not just simulate the design?"
Show the solution
"STA checks every path in the design against the clock, using delays alone, with no test patterns. Simulation only checks the paths a test happens to exercise. A 32-bit adder alone has far too many input patterns to try. So STA gives complete timing coverage, while simulation is still needed to check what the design does."
What is slack?
"What is slack?"
Show the solution
"The margin a timing check passes by. For setup it is the required time minus the arrival time; for hold it is the arrival time minus the required time. Positive slack passes, negative fails, and its size says by how much."
Setup and hold
"What are setup and hold, and which one does the clock period affect?"
Show the solution
"Setup says the data must arrive a little before the capturing edge. Hold says it must stay steady a little after it, so new data must not arrive too soon. The period is in the setup check, because the capture edge is one period after the launch edge. Hold compares data against the same edge, so the period drops out."
Why not slow the clock?
"A chip comes back with a hold violation. Can you fix it by running the clock slower?"
Show the solution
"No. Hold compares the new data with the capture made at the same edge that launched it, so the period never appears in the check. The chip fails at every frequency. That is why hold is fixed with care before tape-out: once it is in silicon, nothing on the board can fix it."
Is skew good or bad?
"Is clock skew good or bad?"
Show the solution
"Neither, on its own. If the capture clock arrives later than the launch clock, setup gets that extra time and hold loses it. Designers use this on purpose as useful skew, lending time to a tight stage from an easy one. What matters is that both checks still pass on every path the flip-flop touches."
Ideal or propagated?
"When do you use an ideal clock, and when a propagated one?"
Show the solution
"Before clock tree synthesis there is no tree to measure, so the clock is ideal: it reaches every flip-flop at once, with the skew estimated inside the uncertainty. After CTS, the clock is propagated through the real buffers and wires, and the real skew replaces the estimate. Sign-off always uses propagated clocks."
What is uncertainty made of?
"What goes into clock uncertainty?"
Show the solution
"For setup: the clock's jitter, some margin, and before the clock tree exists, an estimate of the skew. After CTS the skew estimate comes out, because the real skew is now in the report. For hold, jitter usually drops out, because both events happen at the same edge."
False or multicycle?
"What is the difference between a false path and a multicycle path?"
Show the solution
"A false path can never carry data that matters, so it is not timed at all - for example, a path through two multiplexers that can never both select it. A multicycle path is real, but the design only reads its result every N cycles, so it is timed against N cycles instead of one. Get a false path wrong and a real failure goes unchecked."
Why the hold multicycle?
"Why does set_multicycle_path -setup 3 usually need a -hold 2 beside it?"
Show the solution
"Moving the setup check to the third edge also moves the default hold check, to one edge before it - the
second. That demands the data take at least two cycles, which a path of any sensible length fails.
-hold 2 moves the hold check back to the launch edge, where it belongs."
Timing a crossing
"How do you constrain a signal that crosses between two unrelated clocks?"
Show the solution
"The two clocks have no fixed phase, so a normal setup check between them means nothing. I declare them
asynchronous with set_clock_groups, and put a synchroniser on every crossing. Where the crossing
carries a bus, such as a Gray-coded FIFO pointer, I bound it with set_max_delay -datapath_only so its
bits stay close together."
Time borrowing
"What is time borrowing?"
Show the solution
"A latch is transparent for part of the cycle. If data arrives after the latch opens but before it closes, the latch still passes it on - the stage has borrowed time from the next one. It lets uneven logic share the cycle. The limit is the time the latch is open, less its setup time."
Virtual clocks
"What is a virtual clock for?"
Show the solution
"It describes a clock that exists outside the chip, such as the clock of the chip that sends us data. It has no source pin inside our design. Input and output delays are given against it, so the tool knows when outside data really leaves or must arrive, even when that clock is shifted from ours."
Which corner?
"At which corner do you check setup, and at which hold?"
Show the solution
"Setup is usually worst at the slow corner - slow transistors, low voltage, high temperature - and hold at the fast corner. But sign-off checks both at every corner. Wires and cells do not scale together, so a path's worst case is not always where you expect it."
OCV
"What is on-chip variation, and why derate?"
Show the solution
"Two identical cells on the same chip do not have exactly the same delay. OCV derating models this by making one side of each check slow and the other fast: for setup, the launch path late and the capture clock early. It turns 'every cell is typical' into a safe worst case within one corner."
CRPR
"What is CRPR?"
Show the solution
"Clock reconvergence pessimism removal. Launch and capture clocks usually share their first stretch of tree. Derating counts that shared part late on one side and early on the other, but one wire cannot be both at once. CRPR gives that difference back as a credit in the report."
Crosstalk
"How does crosstalk affect setup and hold?"
Show the solution
"A neighbouring wire couples to a net through the capacitance between them. If it switches the opposite way, the net is slowed, which hurts setup. If it switches the same way, the net is sped up, which hurts hold. A quiet net can also get a glitch. Spacing, shielding and stronger drivers reduce it."
Input delay, max and min
"What do set_input_delay -max and -min mean?"
Show the solution
"They say when data from outside arrives at the input pin, measured from a clock edge. The -max value is the
latest arrival: the sender's slowest clock-to-out plus the slowest trace. Setup uses it. The -min
value is the earliest, with both at their fastest, and hold uses it."
Unconstrained paths
"What happens to a path that has no constraint?"
Show the solution
"Nothing - and that is the danger. An input with no input delay, or an output with no output delay, is simply not checked. It will not show as a violation, so the report looks clean. That is why every sign-off flow runs a check for unconstrained paths."
Fixing hold
"How do you fix a hold violation, and when in the flow?"
Show the solution
"Add delay to the short data path, usually with delay cells, placed on the short branch only. Or reduce the skew that caused it. It is done after clock tree synthesis, because hold depends on real skew. Every cell added must be re-checked for setup at the slow corner."
WNS and TNS
"What do WNS and TNS tell you?"
Show the solution
"WNS, the worst negative slack, says how far the single worst path is from passing. TNS, the total negative slack, adds up every failing path. A design with a bad WNS but a small TNS has one hard path. A small WNS with a large TNS has many near-misses, which usually means a wider problem."
Which of these can fix a hold violation?
Show the answer
Answer: C. Hold needs the data to arrive later, so it needs delay on the data path. The period is not in the hold check, and a multicycle setup makes hold worse. A faster flip-flop sends the data even earlier.
15.4 Report-reading drills
A timing report tells you what is wrong, if you read it in the right order. Read the slack, then the two clock lines, then the biggest step in the Incr column, then the clock edges.
Each report below was printed by the same timing model as the rest of the course, in the layout sign-off tools use. Find the problem before you open the answer.
- Slack - how bad is it? 2. The two clock network delay lines - is skew the problem? 3. The biggest Incr - which cell or wire takes the time? 4. The clock edges - is this the check you meant?
Drill 1
Startpoint: r_acc[3]
Endpoint: r_sum[7]
Path Group: clk
Path Type: max
Point Incr Path
----------------------------------------------------------------
clock clk (rise edge) 0.00 0.00
clock source latency 0.40 0.40
clock network delay (propagated) 1.15 1.55
r_acc[3]/CK 0.00 1.55
r_acc[3]/Q (clock-to-Q) 0.18 1.73
u1/Y (NAND2_X1) 0.14 1.87
u2/Y (OAI21_X1) 0.22 2.09
u3/S (ADD8_X1) 1.98 4.07
u4/Y (MUX2_X1) 0.21 4.28
r_sum[7]/D 0.00 4.28
data arrival time 4.28
clock clk (rise edge) 3.00 3.00
clock source latency 0.40 3.40
clock network delay (propagated) 0.82 4.22
r_sum[7]/CK 0.00 4.22
clock uncertainty -0.06 4.16
library setup time -0.08 4.08
data required time 4.08
----------------------------------------------------------------
data required time 4.08
data arrival time -4.28
----------------------------------------------------------------
slack (VIOLATED) -0.20
Drill 1: the logic is fine
This path fails setup by 0.20 ns. Before you touch the logic, what is wrong, and what would you change?
Show the solution
Read the two clock network delay lines. The clock reaches r_acc[3] 1.15 ns after the source, but r_sum[7] after only 0.82. The skew is 0.82 - 1.15 = -0.33 ns: the capture clock is early, and that costs 0.33 ns.
The 2.55 ns of logic is not the problem. With the capture clock also at 1.15, the same path passes with 0.13 ns. So the fix belongs in the clock tree. Delaying r_sum[7]'s clock helps this path but costs hold on paths into r_sum[7], so check those too.
Drill 2
Startpoint: r_cnt[0]
Endpoint: r_cmp[0]
Path Group: clk
Path Type: min
Point Incr Path
----------------------------------------------------------------
clock clk (rise edge) 0.00 0.00
clock source latency 0.40 0.40
clock network delay (propagated) 0.80 1.20
r_cnt[0]/CK 0.00 1.20
r_cnt[0]/Q (clock-to-Q) 0.10 1.30
u9/Y (BUF_X1) 0.06 1.36
r_cmp[0]/D 0.00 1.36
data arrival time 1.36
clock clk (rise edge) 0.00 0.00
clock source latency 0.40 0.40
clock network delay (propagated) 0.98 1.38
r_cmp[0]/CK 0.00 1.38
clock uncertainty 0.03 1.41
library hold time 0.05 1.46
data required time 1.46
----------------------------------------------------------------
data arrival time 1.36
data required time -1.46
----------------------------------------------------------------
slack (VIOLATED) -0.10
Drill 2: how many cells?
This path fails hold. Why? How many delay cells of 0.03 ns (0.05 ns at the slow corner) fix it, if the path has 2.77 ns of setup slack?
Show the solution
Skew again, the other way: the capture clock arrives at 0.98, 0.18 ns after the launch clock. The data arrives at 1.36 but must not come before 1.46, so the slack is -0.10 ns.
0.10 / 0.03 = 3.33, so 4 cells, giving hold slack 0.02 ns. Setup falls from 2.77 to 2.57 ns - plenty left.
Drill 3
Startpoint: r_op[1]
Endpoint: r_res[4]
Path Group: clk
Path Type: max
Point Incr Path
----------------------------------------------------------------
clock clk (rise edge) 0.00 0.00
clock source latency 0.30 0.30
clock network delay (ideal) 0.00 0.30
r_op[1]/CK 0.00 0.30
r_op[1]/Q (clock-to-Q) 0.14 0.44
u7/Y (AOI22_X1) 0.19 0.63
u8/Z (ALU_X1) 1.92 2.55
r_res[4]/D 0.00 2.55
data arrival time 2.55
clock clk (rise edge) 2.50 2.50
clock source latency 0.30 2.80
clock network delay (ideal) 0.00 2.80
r_res[4]/CK 0.00 2.80
clock uncertainty -0.25 2.55
library setup time -0.08 2.47
data required time 2.47
----------------------------------------------------------------
data required time 2.47
data arrival time -2.55
----------------------------------------------------------------
slack (VIOLATED) -0.08
Drill 3: fix it now?
This report comes from before clock tree synthesis, and fails by 0.08 ns. The 0.25 ns of uncertainty is 0.05 of jitter, a skew estimate of 0.18 and 0.02 of margin. Should you fix it now?
Show the solution
The word ideal on the clock lines says there is no clock tree yet. After CTS, the 0.18 ns guess goes, leaving 0.07 ns, and the real skew takes its place.
With real trees of 0.70 and 0.74 the path passes with 0.14 ns. With 0.74 and 0.62 it fails by 0.02 ns. So this path is worth watching, not rushing. A large failure before CTS is different: no clock tree will rescue it.
Drill 4
Startpoint: r_status[2]
Endpoint: r_view[2]
Path Group: clk_b
Path Type: max
Point Incr Path
----------------------------------------------------------------
clock clk_a (rise edge) 30.00 30.00
r_status[2]/CK 0.00 30.00
r_status[2]/Q (clock-to-Q) 0.20 30.20
u12/Y (XOR2_X1) 0.35 30.55
u13/Y (NOR3_X1) 0.28 30.83
u14/Y (AO22_X1) 1.32 32.15
r_view[2]/D 0.00 32.15
data arrival time 32.15
clock clk_b (rise edge) 32.00 32.00
r_view[2]/CK 0.00 32.00
library setup time -0.08 31.92
data required time 31.92
----------------------------------------------------------------
data required time 31.92
data arrival time -32.15
----------------------------------------------------------------
slack (VIOLATED) -0.23
Drill 4: only two nanoseconds
Only 1.95 ns of logic, and it fails. Look at the clock edges. What is going on? What would you do if clk_a and clk_b come from separate oscillators?
Show the solution
The launch edge is clk_a at 30 and the capture edge clk_b at 32. The tool assumed the two clocks are related and found their tightest pair: 2 ns, the GCD of 10 and 8. Inside one 10 ns clock, the same path would have 7.77 ns to spare.
If the clocks come from separate oscillators, that 2 ns is fiction: their edges drift past each other,
and any gap can occur. Declare them asynchronous with set_clock_groups, and make sure the signal
goes through a synchroniser, as in
Verilog Volume 05. If they do
come from one PLL, 2 ns is the real requirement and the logic must fit it.
Drill 5
Startpoint: r_x[5]
Endpoint: r_y[5]
Path Group: clk
Path Type: max
Point Incr Path
----------------------------------------------------------------
clock clk (rise edge) 0.00 0.00
clock network delay (propagated) 0.94 0.94
r_x[5]/CK 0.00 0.94
r_x[5]/Q (clock-to-Q) 0.19 1.13
u20/Y (NAND3_X1) 0.25 1.39
u21/CO (FA_X1) 3.41 4.80
r_y[5]/D 0.00 4.80
data arrival time 4.80
clock clk (rise edge) 4.00 4.00
clock network delay (propagated) 0.90 4.90
r_y[5]/CK 0.00 4.90
clock reconvergence pessimism 0.07 4.97
clock uncertainty -0.06 4.91
library setup time -0.08 4.83
data required time 4.83
----------------------------------------------------------------
data required time 4.83
data arrival time -4.80
----------------------------------------------------------------
slack (MET) 0.03
Drill 5: the small line
This path passes by 0.03 ns. What is the 0.07 ns line, and would the path pass without it? Why do the Incr values not quite add up to the Path column?
Show the solution
It is CRPR. The first 0.70 ns of both clock paths is the same wire. Derated by 5%, it was counted 0.735 late for launch and 0.665 early for capture. The credit is the difference, 0.07 ns. Without it the path would fail by 0.04 ns - a failure that cannot happen on silicon.
The Incr values are derated and then rounded: 0.90 x 1.05 = 0.945 is printed as 0.94. The Path column adds the unrounded numbers, so the rounded steps can be 0.01 ns off. Trust the Path column.
In a setup report, the launch clock network delay is 1.20 ns and the capture clock network delay is 0.90 ns. What is the skew, and what does it do?
Show the answer
Answer: B. Skew is capture minus launch: 0.90 - 1.20 = -0.30 ns. The capture edge arrives early, so the data has 0.30 ns less time. The 2.10 comes from adding the two instead of subtracting.
15.5 Flashcards
A few minutes of flashcards each day beats a whole night of reading. Say the answer first, then check it.
Click a card to see its answer, and click again to hide it. Try to say the answer out loud first: the effort of remembering is what makes it stick.
01 Slack
02 Setup time
03 Hold time
04 Clock-to-Q
05 Setup slack in one line
06 Hold slack in one line
07 Why does hold not depend on the clock period?
08 f_max
09 Skew
10 Positive skew
11 Useful skew
12 Jitter
13 Setup uncertainty before CTS
14 Ideal clock
15 Propagated clock
16 Source latency
17 Network latency
18 Half-cycle path
19 Two related clocks
20 Unrelated clocks
21 Generated clock
22 Clock-gating check
23 False path
24 -setup N
25 -hold N - 1
26 -start and -end
27 set_max_delay -datapath_only
28 Case analysis
29 Latch
30 Time borrowing
31 Most a latch can lend
32 set_input_delay -max
33 set_output_delay -min
34 Virtual clock
35 Source-synchronous
36 DDR
37 PVT corner
38 RC corner
39 OCV derating
40 CRPR
41 AOCV
42 POCV
43 Crosstalk delta delay
44 Glitch bound
45 Shielding
46 Unconstrained path
47 Delay cell
48 ECO
49 WNS
50 TNS
What does CRPR remove from a timing report?
Show the answer
Answer: A. Derating makes the launch clock late and the capture clock early. The part of the tree they share is one wire, and cannot be both. CRPR gives back that impossible difference. It does not touch real skew.
What you learned in this course
- Explain what static timing analysis checks, and why it replaced simulation for timing.
- Trace where delay comes from: gates, wires, slew and library tables.
- Work out setup and hold slack by hand, with clock trees, skew and uncertainty.
- Handle half-cycle paths, related and unrelated clocks, generated clocks and clock gating.
- Write false paths, multicycles and max delays, and know which checks each one moves.
- Time latches, chip-to-chip interfaces and DDR buses.
- Account for corners, OCV, CRPR, AOCV, POCV and crosstalk.
- Write complete SDC, read any timing report, and close timing with WNS and TNS as your guide.
Key words from this volume
Every word below has a plain-English entry in the glossary.
- Static timing analysis (STA)
- On-chip variation (OCV)
- CRPR
- AOCV
- POCV
- WNS (worst negative slack)
- TNS (total negative slack)
Where to go next
Every number in this course was produced by a timing model and checked before it reached a page. The same habit will serve you in real work: never trust a slack you have not traced line by line.
Here is where you can use what you know next on BlinkNBuild:
- Verilog Volume 07 covers the same ground in one volume, from the RTL designer's side - a quick refresher.
- FPGA Volume 03 applies these constraints in Vivado, as XDC.
- ASIC Volume 06 runs multi-corner sign-off with OpenSTA.
- State Machines Volume 13 times the loop at the heart of every state machine.
- The learning paths suggest what to take next, and the GATE ECE digital circuits questions give you more exam-style practice.
Congratulations on finishing Static Timing Analysis from Zero.