Volume 14 Advanced 5 sub-modules ~20 min read

Timing Closure

Everything so far has been about reading timing. This volume is about changing it. When the slack is negative there are a dozen things you can do, and they differ hugely in how much they buy and what they cost. Here is each one, measured on a real path, and then a small design taken from its first failing report all the way to sign-off.

You will learn
  • The ways to fix a setup violation, and how much each one buys
  • How to fix hold with delay cells without breaking setup
  • What an ECO is, and why late fixes are made so carefully
  • Why timing closure on an FPGA uses different tools from an ASIC
  • How a real closure goes, step by step, with WNS and TNS
You need
  • Volumes 04 to 06 for setup, hold and the clock, and Volume 11 for corners

14.1 Fixing setup violations

A setup violation is fixed in one of two ways: make the path faster, or give it more time. The fixes differ hugely in how much they buy and what they cost, so the first step is always to read where the time goes.

Timing closure is the work of changing a design until every path passes. This volume uses the adder path again, now at 250 MHz (4 ns), with slow-corner delays, since that is where setup is signed off.

The adder path at 250 MHz with slow-corner delays, failing setup u_ff1 D Q t_cq 0.21 NAND2 0.18 ns XOR2 0.31 ns 16-bit adder 2.60 ns AOI21 0.26 ns MUX2_X1 0.42 ns setup 0.09 u_ff2 D Q clk period 4.00 ns
Figure 14.1 - The same five cells as Volume 04, at the slow corner. The data takes 0.21 + 3.77 = 3.98 ns. A 4 ns clock, less 0.09 ns of setup and 0.08 ns of uncertainty, only allows 3.83 ns. The path fails by 0.15 ns.

Read the Incr column first. The adder is 2.60 of the 3.77 ns of logic. The multiplexer, 0.42 ns, is slow for such a small cell - a sign it is too weak for its load.

Every fix, measured

Fix Slack Change
As built -0.15 ns -
Upsize the multiplexer, X1 to X2 (0.42 to 0.28 ns) -0.01 ns +0.14
Low-Vt cells in the adder (2.60 to 2.34 ns) +0.11 ns +0.26
A faster adder architecture (2.60 to 1.95 ns) +0.50 ns +0.65
Useful skew: capture clock 0.15 ns later 0.00 ns +0.15
Upsize and low-Vt together +0.25 ns +0.40
Pipeline after the adder +0.53 and +2.94 ns one more cycle of latency
  1. Upsizing is cheap and local, but buys little: a stronger gate only helps where the load was the problem.
  2. Low-Vt cells are faster versions of the same cells. They leak more power while idle, so they go only where needed.
  3. A different architecture buys the most, but it is a design change, made early or not at all.
  4. Useful skew moves time from the next stage, as in Volume 06; that stage must have it to spare.
  5. Pipelining splits the path, at the cost of a cycle of latency the design must be able to accept.
Remember

Fix the biggest piece of delay with the biggest lever, then tidy up with the small ones. Upsizing a 0.18 ns gate cannot rescue a path whose adder is 2.60 ns.

Common mistake

Reaching for useful skew first because it needs no logic change. It only moves slack between neighbouring stages. On a design where most paths are tight, it just moves the violation next door.

Quick check

A path fails by 0.30 ns. One cell on it takes 1.80 ns of its 2.40 ns of logic. What should you look at first?

Show the answer

Answer: B. Three-quarters of the delay is in one cell, so that is where 0.30 ns is easiest to find. Small gates could not give back that much. Delay cells would make setup worse; they are a hold fix.

14.2 Fixing hold violations

Hold violations are fixed by adding delay to the short path, usually with delay cells. Every cell added for hold also slows the path for setup, so each fix must be checked at the slow corner too.

A delay cell adds delay at the fast corner, where hold is checked. It also adds more delay at the slow corner, where setup is checked. Take delay cells of 0.04 ns at the fast corner and 0.06 ns at the slow one:

Path Hold slack Setup slack Cells needed Hold after Setup after
P1 -0.08 ns 1.20 ns 2 0.00 ns 1.08 ns
P2 -0.05 ns 0.60 ns 2 0.03 ns 0.48 ns
P3 -0.03 ns 0.05 ns 1 0.01 ns -0.01 ns
P4 -0.01 ns 2.00 ns 1 0.03 ns 1.94 ns

The cells needed are the hold violation divided by 0.04 ns, rounded up. Three fixes are clean. P3 is the trap: its hold is fixed, but its setup now fails.

When a hold fix breaks setup

P3 has a short path and a long path meeting at the same capture flip-flop - one sets its hold, the other its setup. A delay cell at the flip-flop's input slows both. The answer is to put the delay only on the short branch, before the paths join, or to fix the skew instead of the data path.

In plain words

Hold and setup pull the same path in opposite directions. Hold wants it slower at the fast corner; setup wants it faster at the slow corner. A hold fix is only a fix if both checks still pass.

Common mistake

Fixing hold before the clock tree is built. With an ideal clock and a guessed uncertainty, the tool pads paths that were never going to fail, and every padded path is slower for setup. Fix hold once the real clock tree exists.

Quick check

A path fails hold by 0.10 ns. Delay cells add 0.04 ns at the fast corner and 0.06 ns at the slow corner. How many are needed, and how much setup slack do they cost?

Show the answer

Answer: C. 0.10 / 0.04 = 2.5, rounded up to 3 cells. They add 3 x 0.04 = 0.12 ns for hold, and 3 x 0.06 = 0.18 ns of setup delay at the slow corner.

14.3 Engineering change orders (ECOs)

An ECO is a small, targeted change to a nearly finished design - resizing a few cells, adding a few delay cells - made without running the whole flow again. The later it comes, the smaller and more careful it must be.

Near the end of a project, running synthesis and place-and-route again would change thousands of cells, and every timing result with them. So the last fixes are made as ECOs: a list of exact changes, applied in place, with the rest of the design left untouched.

ECO step Example
Identify The report shows P3 failing setup after a hold fix
Change Move P3's delay cell to the short branch; upsize one gate on the long branch
Legalise Place the changed cells in free space without moving their neighbours
Re-route Re-route only the nets that changed
Re-time Run timing again at every corner, not just the one that failed

Metal-only ECOs

After the masks for the transistors are made, a change can still be made by altering only the metal wiring layers, which is far cheaper. It works because designers scatter spare cells - unused gates - across the chip in advance. A metal-only ECO connects them in.

Remember

Every ECO is re-timed at every corner and mode. Most late surprises are an ECO that fixed one corner and broke another.

Quick check

Why do chips carry spare cells that are not connected to anything?

Show the answer

Answer: A. Changing the transistor layers after masks are made is very expensive. Spare cells already exist on the chip, so a fix can wire them in by changing only the cheaper metal layers.

14.4 Closure on FPGA versus ASIC

The arithmetic of closure is the same on an FPGA and an ASIC, but the levers are not. An FPGA's cells are fixed and its wires are slow, so closure there is mostly about placement, routing and pipelining.

On an ASIC you choose every cell's size and threshold voltage. On an FPGA the logic cells are fixed, and most of a path's delay is routing between them. Take a path at 300 MHz (3.333 ns), with five look-up tables of 0.12 ns and five routes of 0.60 ns:

Change Slack
As placed -0.68 ns (routing is 83% of the data path)
Placed tighter, routes of 0.40 ns +0.32 ns
Pipelined into stages of 3 and 2 look-up tables +0.76 and +1.48 ns
Lever ASIC FPGA
Cell size and threshold voltage Yes No - the cells are fixed
Restructure logic Yes Yes, through synthesis settings
Placement Yes Yes - often the biggest lever
Pipelining and retiming Yes Yes - very effective
Delay cells for hold Yes No - the router fixes hold by adding routing delay
Useful skew Yes Limited: clock networks are fixed
In plain words

On an ASIC you fix a slow path by building better gates. On an FPGA you fix it by using fewer, closer ones - or by giving it another clock cycle.

The FPGA tools that do this work are covered hands-on in FPGA Mastery Volume 02. For the ASIC side, see the clock tree and useful skew in ASIC Volume 04, and the sign-off loop in ASIC Volume 06.

Quick check

An FPGA path fails setup, and most of its delay is routing. Which fix is not available?

Show the answer

Answer: D. FPGA logic cells come in one size, fixed in the silicon. Placement, pipelining and restructuring all work; upsizing does not exist on an FPGA.

14.5 A full closure walkthrough

Real closure is a loop: time the design, fix the worst problems, and time it again - first for setup, then, once the clock tree exists, for hold. WNS and TNS track the progress.

A small design of five paths at 250 MHz. Setup is checked at the slow corner and hold at the fast one. WNS is the worst slack; TNS adds up every failing slack.

Step 1 - after placement, with an ideal clock

Path Setup slack Hold slack
A (the adder) -0.25 ns 1.42 ns
B (a multiplier) -0.14 ns 1.18 ns
C (a decoder) 0.32 ns 0.88 ns
D (one buffer) 3.32 ns 0.04 ns
E (one inverter) 3.38 ns 0.06 ns

Setup WNS -0.25 ns, TNS -0.39 ns, two paths failing. The uncertainty is still the early estimate: 0.18 ns for setup and 0.13 ns for hold. Hold looks clean - for now.

Step 2 - setup fixes

Path A gets its multiplexer upsized and its adder in low-Vt cells. Path B's multiplier is restructured, from 3.66 to 3.40 ns.

Path Setup slack Hold slack
A 0.15 ns 1.42 ns
B 0.12 ns 1.08 ns
C 0.32 ns 0.88 ns
D 3.32 ns 0.04 ns
E 3.38 ns 0.06 ns

Setup WNS +0.12 ns. Every setup check passes. Notice that B's hold slack dropped from 1.18 to 1.08 ns: the restructured logic is faster at both corners.

Step 3 - after clock tree synthesis

The clock is now propagated through a real tree, and the uncertainty drops to 0.08 ns and 0.03 ns.

Path Setup slack Hold slack
A 0.28 ns 1.49 ns
B 0.16 ns 1.24 ns
C 0.48 ns 0.92 ns
D 3.61 ns -0.05 ns
E 3.65 ns -0.01 ns

Setup improved, because the pessimistic early uncertainty is gone. But hold now fails on D and E. The real tree reaches their capture flip-flops 0.19 and 0.17 ns late - skew no estimate predicted. Hold WNS -0.05 ns, TNS -0.06 ns.

Step 4 - hold fixes

D gets two delay cells, E gets one.

Path Setup slack Hold slack
A 0.28 ns 1.49 ns
B 0.16 ns 1.24 ns
C 0.48 ns 0.92 ns
D 3.49 ns 0.03 ns
E 3.59 ns 0.03 ns

Every check passes: setup WNS +0.16 ns, hold WNS +0.03 ns. D and E had setup slack to burn, so their delay cells cost nothing that mattered.

Setup and hold worst negative slack through the four closure steps 1 2 3 4 -0.3 -0.2 -0.1 0 0.1 0.2 closure step worst slack (ns) setup WNS hold WNS
Figure 14.2 - Setup (gold) is fixed in step 2 and improves again in step 3, when the early uncertainty is replaced by the real tree. Hold (cyan) looks fine until the real clock tree arrives in step 3, then is fixed with delay cells in step 4. Below the dashed zero line, a check fails.
Remember

This order is not an accident: fix setup with an ideal clock, build the clock tree, then fix hold. Fixing hold first would pad paths against skew that was only guessed.

Quick check

A design has setup slacks of -0.30, -0.12, -0.05 and +0.40 ns. What are its WNS and TNS?

Show the answer

Answer: B. WNS is the worst single slack: -0.30 ns. TNS adds only the failing ones: -0.30 - 0.12 - 0.05 = -0.47 ns. The passing +0.40 ns path is not counted.

What you learned

Key words from this volume

Every word below has a plain-English entry in the glossary.

Practice

Practice 1

Choose the fixes

The adder path fails by 0.15 ns. You may use at most two fixes, and the design cannot accept another cycle of latency. Which two, and what slack do you end up with?

Show the solution

Upsize the multiplexer and use low-Vt cells in the adder. Together they give +0.25 ns, from the table in 14.1.

The faster adder architecture alone would give +0.50 ns, but it is a design change; if the RTL is frozen, it is not on the table. Pipelining is ruled out by the latency.

Practice 2

Useful skew, checked

Useful skew of 0.15 ns on the adder path's capture flip-flop brings its setup slack to exactly 0.00 ns. Is that a good fix?

Show the solution

Not on its own. Zero slack passes, with no margin, and the same 0.15 ns comes out of the next stage's setup and this flip-flop's hold. It is only a good fix if the next stage has slack to spare and the hold check still passes. Even then, combine it with something that adds real margin, such as the upsize.

Interview corner

Interview question 1

Your design fails setup by 200 ps

"Your worst setup slack is -200 ps across 30 paths. How do you close it?"

Show the solution

"First I would look at what the 30 paths have in common - often one cell, one net or one block. Then I would read the worst report line by line, to see whether the delay is one big cell, lots of small ones, or wire. A big cell calls for a faster architecture, low-Vt or upsizing; lots of logic levels call for restructuring or pipelining; long wires call for placement or buffering. I would check the constraints too, since a missing multicycle can look like 30 failing paths. After each change I would re-time at every corner, watching TNS as well as WNS."

Interview question 2

Why fix hold after CTS?

"Why is hold usually fixed after clock tree synthesis, not before?"

Show the solution

"Because hold depends on skew, and skew is only known once the clock tree exists. Before CTS the tool uses an ideal clock with a guessed uncertainty, so it would pad paths that turn out fine and miss ones that turn bad. Every delay cell added for hold also costs setup, so padding the wrong paths hurts. After CTS, with propagated clocks, the hold picture is real, and the tool adds the smallest delay that clears each violation."

Volume 15 gathers the whole course into one place for revision: a formula sheet, forty numerical problems, twenty concept questions, report-reading drills and flashcards.