Timing Closure
Everything so far has been about reading timing. This volume is about changing it. When the slack is negative there are a dozen things you can do, and they differ hugely in how much they buy and what they cost. Here is each one, measured on a real path, and then a small design taken from its first failing report all the way to sign-off.
- The ways to fix a setup violation, and how much each one buys
- How to fix hold with delay cells without breaking setup
- What an ECO is, and why late fixes are made so carefully
- Why timing closure on an FPGA uses different tools from an ASIC
- How a real closure goes, step by step, with WNS and TNS
- Volumes 04 to 06 for setup, hold and the clock, and Volume 11 for corners
14.1 Fixing setup violations
A setup violation is fixed in one of two ways: make the path faster, or give it more time. The fixes differ hugely in how much they buy and what they cost, so the first step is always to read where the time goes.
Timing closure is the work of changing a design until every path passes. This volume uses the adder path again, now at 250 MHz (4 ns), with slow-corner delays, since that is where setup is signed off.
Read the Incr column first. The adder is 2.60 of the 3.77 ns of logic. The multiplexer, 0.42 ns, is slow for such a small cell - a sign it is too weak for its load.
Every fix, measured
| Fix | Slack | Change |
|---|---|---|
| As built | -0.15 ns | - |
| Upsize the multiplexer, X1 to X2 (0.42 to 0.28 ns) | -0.01 ns | +0.14 |
| Low-Vt cells in the adder (2.60 to 2.34 ns) | +0.11 ns | +0.26 |
| A faster adder architecture (2.60 to 1.95 ns) | +0.50 ns | +0.65 |
| Useful skew: capture clock 0.15 ns later | 0.00 ns | +0.15 |
| Upsize and low-Vt together | +0.25 ns | +0.40 |
| Pipeline after the adder | +0.53 and +2.94 ns | one more cycle of latency |
- Upsizing is cheap and local, but buys little: a stronger gate only helps where the load was the problem.
- Low-Vt cells are faster versions of the same cells. They leak more power while idle, so they go only where needed.
- A different architecture buys the most, but it is a design change, made early or not at all.
- Useful skew moves time from the next stage, as in Volume 06; that stage must have it to spare.
- Pipelining splits the path, at the cost of a cycle of latency the design must be able to accept.
Fix the biggest piece of delay with the biggest lever, then tidy up with the small ones. Upsizing a 0.18 ns gate cannot rescue a path whose adder is 2.60 ns.
Reaching for useful skew first because it needs no logic change. It only moves slack between neighbouring stages. On a design where most paths are tight, it just moves the violation next door.
A path fails by 0.30 ns. One cell on it takes 1.80 ns of its 2.40 ns of logic. What should you look at first?
Show the answer
Answer: B. Three-quarters of the delay is in one cell, so that is where 0.30 ns is easiest to find. Small gates could not give back that much. Delay cells would make setup worse; they are a hold fix.
14.2 Fixing hold violations
Hold violations are fixed by adding delay to the short path, usually with delay cells. Every cell added for hold also slows the path for setup, so each fix must be checked at the slow corner too.
A delay cell adds delay at the fast corner, where hold is checked. It also adds more delay at the slow corner, where setup is checked. Take delay cells of 0.04 ns at the fast corner and 0.06 ns at the slow one:
| Path | Hold slack | Setup slack | Cells needed | Hold after | Setup after |
|---|---|---|---|---|---|
| P1 | -0.08 ns | 1.20 ns | 2 | 0.00 ns | 1.08 ns |
| P2 | -0.05 ns | 0.60 ns | 2 | 0.03 ns | 0.48 ns |
| P3 | -0.03 ns | 0.05 ns | 1 | 0.01 ns | -0.01 ns |
| P4 | -0.01 ns | 2.00 ns | 1 | 0.03 ns | 1.94 ns |
The cells needed are the hold violation divided by 0.04 ns, rounded up. Three fixes are clean. P3 is the trap: its hold is fixed, but its setup now fails.
When a hold fix breaks setup
P3 has a short path and a long path meeting at the same capture flip-flop - one sets its hold, the other its setup. A delay cell at the flip-flop's input slows both. The answer is to put the delay only on the short branch, before the paths join, or to fix the skew instead of the data path.
Hold and setup pull the same path in opposite directions. Hold wants it slower at the fast corner; setup wants it faster at the slow corner. A hold fix is only a fix if both checks still pass.
Fixing hold before the clock tree is built. With an ideal clock and a guessed uncertainty, the tool pads paths that were never going to fail, and every padded path is slower for setup. Fix hold once the real clock tree exists.
A path fails hold by 0.10 ns. Delay cells add 0.04 ns at the fast corner and 0.06 ns at the slow corner. How many are needed, and how much setup slack do they cost?
Show the answer
Answer: C. 0.10 / 0.04 = 2.5, rounded up to 3 cells. They add 3 x 0.04 = 0.12 ns for hold, and 3 x 0.06 = 0.18 ns of setup delay at the slow corner.
14.3 Engineering change orders (ECOs)
An ECO is a small, targeted change to a nearly finished design - resizing a few cells, adding a few delay cells - made without running the whole flow again. The later it comes, the smaller and more careful it must be.
Near the end of a project, running synthesis and place-and-route again would change thousands of cells, and every timing result with them. So the last fixes are made as ECOs: a list of exact changes, applied in place, with the rest of the design left untouched.
| ECO step | Example |
|---|---|
| Identify | The report shows P3 failing setup after a hold fix |
| Change | Move P3's delay cell to the short branch; upsize one gate on the long branch |
| Legalise | Place the changed cells in free space without moving their neighbours |
| Re-route | Re-route only the nets that changed |
| Re-time | Run timing again at every corner, not just the one that failed |
Metal-only ECOs
After the masks for the transistors are made, a change can still be made by altering only the metal wiring layers, which is far cheaper. It works because designers scatter spare cells - unused gates - across the chip in advance. A metal-only ECO connects them in.
Every ECO is re-timed at every corner and mode. Most late surprises are an ECO that fixed one corner and broke another.
Why do chips carry spare cells that are not connected to anything?
Show the answer
Answer: A. Changing the transistor layers after masks are made is very expensive. Spare cells already exist on the chip, so a fix can wire them in by changing only the cheaper metal layers.
14.4 Closure on FPGA versus ASIC
The arithmetic of closure is the same on an FPGA and an ASIC, but the levers are not. An FPGA's cells are fixed and its wires are slow, so closure there is mostly about placement, routing and pipelining.
On an ASIC you choose every cell's size and threshold voltage. On an FPGA the logic cells are fixed, and most of a path's delay is routing between them. Take a path at 300 MHz (3.333 ns), with five look-up tables of 0.12 ns and five routes of 0.60 ns:
| Change | Slack |
|---|---|
| As placed | -0.68 ns (routing is 83% of the data path) |
| Placed tighter, routes of 0.40 ns | +0.32 ns |
| Pipelined into stages of 3 and 2 look-up tables | +0.76 and +1.48 ns |
| Lever | ASIC | FPGA |
|---|---|---|
| Cell size and threshold voltage | Yes | No - the cells are fixed |
| Restructure logic | Yes | Yes, through synthesis settings |
| Placement | Yes | Yes - often the biggest lever |
| Pipelining and retiming | Yes | Yes - very effective |
| Delay cells for hold | Yes | No - the router fixes hold by adding routing delay |
| Useful skew | Yes | Limited: clock networks are fixed |
On an ASIC you fix a slow path by building better gates. On an FPGA you fix it by using fewer, closer ones - or by giving it another clock cycle.
The FPGA tools that do this work are covered hands-on in FPGA Mastery Volume 02. For the ASIC side, see the clock tree and useful skew in ASIC Volume 04, and the sign-off loop in ASIC Volume 06.
An FPGA path fails setup, and most of its delay is routing. Which fix is not available?
Show the answer
Answer: D. FPGA logic cells come in one size, fixed in the silicon. Placement, pipelining and restructuring all work; upsizing does not exist on an FPGA.
14.5 A full closure walkthrough
Real closure is a loop: time the design, fix the worst problems, and time it again - first for setup, then, once the clock tree exists, for hold. WNS and TNS track the progress.
A small design of five paths at 250 MHz. Setup is checked at the slow corner and hold at the fast one. WNS is the worst slack; TNS adds up every failing slack.
Step 1 - after placement, with an ideal clock
| Path | Setup slack | Hold slack |
|---|---|---|
| A (the adder) | -0.25 ns | 1.42 ns |
| B (a multiplier) | -0.14 ns | 1.18 ns |
| C (a decoder) | 0.32 ns | 0.88 ns |
| D (one buffer) | 3.32 ns | 0.04 ns |
| E (one inverter) | 3.38 ns | 0.06 ns |
Setup WNS -0.25 ns, TNS -0.39 ns, two paths failing. The uncertainty is still the early estimate: 0.18 ns for setup and 0.13 ns for hold. Hold looks clean - for now.
Step 2 - setup fixes
Path A gets its multiplexer upsized and its adder in low-Vt cells. Path B's multiplier is restructured, from 3.66 to 3.40 ns.
| Path | Setup slack | Hold slack |
|---|---|---|
| A | 0.15 ns | 1.42 ns |
| B | 0.12 ns | 1.08 ns |
| C | 0.32 ns | 0.88 ns |
| D | 3.32 ns | 0.04 ns |
| E | 3.38 ns | 0.06 ns |
Setup WNS +0.12 ns. Every setup check passes. Notice that B's hold slack dropped from 1.18 to 1.08 ns: the restructured logic is faster at both corners.
Step 3 - after clock tree synthesis
The clock is now propagated through a real tree, and the uncertainty drops to 0.08 ns and 0.03 ns.
| Path | Setup slack | Hold slack |
|---|---|---|
| A | 0.28 ns | 1.49 ns |
| B | 0.16 ns | 1.24 ns |
| C | 0.48 ns | 0.92 ns |
| D | 3.61 ns | -0.05 ns |
| E | 3.65 ns | -0.01 ns |
Setup improved, because the pessimistic early uncertainty is gone. But hold now fails on D and E. The real tree reaches their capture flip-flops 0.19 and 0.17 ns late - skew no estimate predicted. Hold WNS -0.05 ns, TNS -0.06 ns.
Step 4 - hold fixes
D gets two delay cells, E gets one.
| Path | Setup slack | Hold slack |
|---|---|---|
| A | 0.28 ns | 1.49 ns |
| B | 0.16 ns | 1.24 ns |
| C | 0.48 ns | 0.92 ns |
| D | 3.49 ns | 0.03 ns |
| E | 3.59 ns | 0.03 ns |
Every check passes: setup WNS +0.16 ns, hold WNS +0.03 ns. D and E had setup slack to burn, so their delay cells cost nothing that mattered.
This order is not an accident: fix setup with an ideal clock, build the clock tree, then fix hold. Fixing hold first would pad paths against skew that was only guessed.
A design has setup slacks of -0.30, -0.12, -0.05 and +0.40 ns. What are its WNS and TNS?
Show the answer
Answer: B. WNS is the worst single slack: -0.30 ns. TNS adds only the failing ones: -0.30 - 0.12 - 0.05 = -0.47 ns. The passing +0.40 ns path is not counted.
What you learned
- Read where the delay goes before choosing a fix; use the biggest lever on the biggest delay.
- Setup is fixed by faster cells, better structure, useful skew or pipelining.
- Hold is fixed with delay cells, and every one must be re-checked for setup at the slow corner.
- ECOs make small, exact changes late in a project; spare cells allow metal-only fixes.
- FPGA closure relies on placement and pipelining, because its cells are fixed.
- Closure runs in order: setup with an ideal clock, then the clock tree, then hold.
- WNS says how far off the worst path is; TNS says how much work is left.
Key words from this volume
Every word below has a plain-English entry in the glossary.
- Timing closure
- Low-Vt cell
- Delay cell
- ECO (engineering change order)
- Spare cells
- Retiming
- WNS (worst negative slack)
- TNS (total negative slack)
Practice
Choose the fixes
The adder path fails by 0.15 ns. You may use at most two fixes, and the design cannot accept another cycle of latency. Which two, and what slack do you end up with?
Show the solution
Upsize the multiplexer and use low-Vt cells in the adder. Together they give +0.25 ns, from the table in 14.1.
The faster adder architecture alone would give +0.50 ns, but it is a design change; if the RTL is frozen, it is not on the table. Pipelining is ruled out by the latency.
Useful skew, checked
Useful skew of 0.15 ns on the adder path's capture flip-flop brings its setup slack to exactly 0.00 ns. Is that a good fix?
Show the solution
Not on its own. Zero slack passes, with no margin, and the same 0.15 ns comes out of the next stage's setup and this flip-flop's hold. It is only a good fix if the next stage has slack to spare and the hold check still passes. Even then, combine it with something that adds real margin, such as the upsize.
Interview corner
Your design fails setup by 200 ps
"Your worst setup slack is -200 ps across 30 paths. How do you close it?"
Show the solution
"First I would look at what the 30 paths have in common - often one cell, one net or one block. Then I would read the worst report line by line, to see whether the delay is one big cell, lots of small ones, or wire. A big cell calls for a faster architecture, low-Vt or upsizing; lots of logic levels call for restructuring or pipelining; long wires call for placement or buffering. I would check the constraints too, since a missing multicycle can look like 30 failing paths. After each change I would re-time at every corner, watching TNS as well as WNS."
Why fix hold after CTS?
"Why is hold usually fixed after clock tree synthesis, not before?"
Show the solution
"Because hold depends on skew, and skew is only known once the clock tree exists. Before CTS the tool uses an ideal clock with a guessed uncertainty, so it would pad paths that turn out fine and miss ones that turn bad. Every delay cell added for hold also costs setup, so padding the wrong paths hurts. After CTS, with propagated clocks, the hold picture is real, and the tool adds the smallest delay that clears each violation."
Volume 15 gathers the whole course into one place for revision: a formula sheet, forty numerical problems, twenty concept questions, report-reading drills and flashcards.