Where Delay Comes From
So far every delay has been a number handed to you. This volume shows where the numbers come from. A gate is slow because it has to charge a capacitance through a resistance. A wire is slower still, because it is made of both. And a timing tool knows all of it from tables - which you will learn to read by hand.
- Why a gate's delay grows with its load, and what a bigger gate buys you
- Why wire delay grows with the square of the length, and how repeaters fix it
- What slew is, and why a slow edge slows down the next gate
- How to read and interpolate a library delay table by hand
- Where a flip-flop's clock-to-Q, setup and hold numbers really come from
2.1 Gate delay and what changes it
A gate is slow because it has to charge a capacitance through a resistance. More load means more to charge, so more delay. A bigger gate has less resistance, so it charges the same load faster.
Inside every gate are transistors: tiny switches. When the input changes, one of them switches on and starts pushing current into the output wire. That wire, and every gate input hanging off it, has to fill up with charge before the voltage reaches the other side.
A switched-on transistor is not a perfect switch. It has resistance, so the current is limited. A little current filling a big capacitance takes a long time.
The 0.69 comes from how a capacitance charges: it reaches half the supply voltage after 0.69 of the time R x C. With R in kilo-ohms and C in femtofarads, R x C comes out in picoseconds.
One gate, three loads, two sizes
This course uses a small made-up library, with numbers chosen to keep the sums readable. Its NAND gate comes in two sizes. The bigger one, X4, has a quarter of the resistance, and four times the input capacitance.
| Load | Capacitance | NAND2_X1 | NAND2_X4 |
|---|---|---|---|
| Fanout 1 | 2.5 fF | 15.2 ps | 13.3 ps |
| Fanout 4 | 8.0 fF | 26.6 ps | 16.2 ps |
| Fanout 16 | 29.0 fF | 70.3 ps | 27.1 ps |
For the X1: 10 ps + 0.69 x 3.00 kOhm x 29.0 fF = 70.3 ps. For the X4: 12 ps + 0.69 x 0.75 kOhm x 29.0 fF = 27.1 ps. On a light load the two are nearly the same. On a heavy load the big gate is more than twice as fast.
Nothing is free: the big gate's input
A bigger gate has bigger transistors, and so a bigger input capacitance. The gate before it now has more to charge. That is the price of drive strength.
| Two gates in a row, fanout 16 at the end | First gate | Second gate | Total |
|---|---|---|---|
| X1 then X1 | 15.2 ps | 70.3 ps | 85.5 ps |
| X1 then X4 | 24.6 ps | 27.1 ps | 51.6 ps |
The first gate got slower, from 15.2 ps to 24.6 ps. The path still got much faster overall. Choosing gate sizes is this trade-off, made millions of times by the synthesis tool.
What else changes a gate's delay
- Temperature. A hot chip is usually slower.
- Supply voltage. A lower voltage pushes less current, so it is slower.
- Manufacturing. No two chips are identical; some come out slow, some fast.
- The input edge. A slowly rising input makes the gate slower. Sub-module 2.3 explains why.
The first three are gathered into PVT corners. For the fanout-4 X1 in this example library, that gives 18.6 ps at the fast corner, 26.6 ps typically, and 38.6 ps at the slow one. Volume 11 is all about them.
Thinking a bigger gate is always faster. It drives its own load faster, but it loads the gate before it more heavily. Upsize a gate whose load is already small and the path can get slower.
A gate has an own delay of 8 ps and a resistance of 2 kOhm. It drives 10 fF. What is its delay?
Show the answer
Answer: B. Delay = 8 + 0.69 x 2 x 10 = 8 + 13.9 = 21.9 ps. The trap answers forget the 0.69, or forget the gate's own delay.
2.2 Wire delay: resistance and capacitance
A wire has resistance and capacitance along its whole length. Make it twice as long and both double, so its own delay goes up four times. Long wires are slow in a way gates never are.
On a small chip, wires are cheap. On a big one, a signal may cross several millimetres, and the wire costs more time than the gates at either end.
The wire is not one resistance and one capacitance. It is many tiny pieces of each, spread out along its length. The charge has to travel through the resistance of the near end to fill the capacitance of the far end.
Here R_d is the driver, R_w and C_w are the wire, and C_L is the load.
Read it term by term. The driver charges everything. The wire charges itself, and 0.38 is the figure for charge spread along a line. Finally, the wire's resistance slows the charging of the load at the far end.
A wire, measured
Take a wire of 1 ohm and 0.2 fF per micrometre, driven through 1 kOhm into a 5 fF load.
| Length | R wire | C wire | Driver term | Wire term | Load term | Total |
|---|---|---|---|---|---|---|
| 100 um | 0.10 kOhm | 20 fF | 17.3 | 0.8 | 0.3 | 18.4 ps |
| 500 um | 0.50 kOhm | 100 fF | 72.8 | 19.0 | 1.7 | 93.5 ps |
| 1000 um | 1.00 kOhm | 200 fF | 142.1 | 76.0 | 3.5 | 221.6 ps |
| 2000 um | 2.00 kOhm | 400 fF | 280.7 | 304.0 | 6.9 | 591.7 ps |
| 4000 um | 4.00 kOhm | 800 fF | 558.0 | 1216.0 | 13.9 | 1787.8 ps |
Watch the wire term: 76.0, then 304.0, then 1216.0. Each doubling of the length multiplies it by four. At 4 mm it is two-thirds of the total.
RC delay grows as length times length. A wire twice as long is not twice as slow - it is up to four times as slow. Nothing else on a chip behaves like that.
The fix: cut the wire into pieces
If a wire's own delay grows with the square of its length, then several short wires are faster than one long one. So long wires are cut into pieces, with a buffer driving each piece. Those buffers are called repeaters.
The same 4 mm wire, cut into equal pieces, with a buffer of 15 ps, 1.0 kOhm and 5 fF before each:
| Pieces | Each piece | Total delay |
|---|---|---|
| 1 | 4000 um | 1802.8 ps |
| 2 | 2000 um | 1213.3 ps |
| 4 | 1000 um | 946.2 ps |
| 8 | 500 um | 868.1 ps |
| 12 | 333 um | 891.3 ps |
| 16 | 250 um | 939.8 ps |
| 32 | 125 um | 1197.3 ps |
Eight pieces bring 1802.8 ps down to 868.1 ps, less than half. Past that, the buffers' own delay starts to win, and more pieces make it worse again. There is always a best number.
Adding repeaters to every wire. On a short wire the wire term is tiny - 0.8 ps at 100 um - so a repeater only adds its own delay. Repeaters pay only on wires long enough for the square law to bite.
A wire's own term is 19.0 ps at 500 um. Nothing else changes. What is it at 1500 um?
Show the answer
Answer: C. The wire term grows with the square of the length. Three times the length gives nine times the term: 9 x 19.0 = 171.0 ps. The first answer grows it in a straight line, which is exactly the mistake the square law punishes.
2.3 Slew, the transition time
Slew is how long an edge takes to get from low to high. A gate with a heavy load makes a slow edge, and a slow edge arriving at the next gate makes that gate slow too. Delay is passed down the path.
So far every edge has been drawn as a straight vertical line. Real edges are ramps. When a gate charges a big capacitance, its output climbs slowly.
The slew of an edge charging through R is about 2.2 x R x C, measured from 10% to 90%.
| Gate | Load | C | Output slew |
|---|---|---|---|
| NAND2_X1 | fanout 1 | 2.5 fF | 16.5 ps |
| NAND2_X1 | fanout 4 | 8.0 fF | 52.7 ps |
| NAND2_X1 | fanout 16 | 29.0 fF | 191.2 ps |
| NAND2_X4 | fanout 16 | 29.0 fF | 47.8 ps |
Why a slow edge slows the next gate
The next gate does not switch the instant its input starts to move. Its transistors only turn fully on once the input is well past halfway. A slow ramp spends a long time getting there, so the next gate starts late - and, part-way on, it switches weakly.
That means delay is not just a property of each gate on its own. A weak gate driving a big load hurts itself and the gate after it. Sub-module 2.4 puts a number on this.
A timing tool carries two numbers along every path: the arrival time, and the slew of that arrival. Each gate's delay depends on the slew coming in, and each gate produces a new slew going out.
Fixing a slow gate by looking only at that gate. If its input edge is slow, the real cure is often upstream: a stronger gate before it, or less load on the net feeding it.
A gate with 1.5 kOhm of resistance drives 20 fF. Roughly what is its output slew?
Show the answer
Answer: A. Slew is about 2.2 x R x C = 2.2 x 1.5 x 20 = 65.9 ps, so about 66 ps. The 21 ps answer uses 0.69 instead of 2.2 - that is the delay factor, not the slew factor.
2.4 Library delay tables made simple
A timing library stores each cell's delay as a small table, looked up by input slew and output load. Between the table's rows and columns, the tool interpolates in straight lines.
The formula in 2.1 gives the right shape, but real transistors do not follow it exactly. So the people who make a cell library measure every cell, at many slews and many loads, and store the results.
The result is an NLDM table: a grid of numbers. Here is our NAND2_X1's delay.
| Input slew \ load | 1 fF | 4 fF | 16 fF | 64 fF |
|---|---|---|---|---|
| 10 ps | 14.1 | 20.4 | 45.4 | 145.5 |
| 50 ps | 22.7 | 29.1 | 54.4 | 155.7 |
| 150 ps | 47.8 | 54.3 | 80.3 | 184.5 |
| 400 ps | 132.3 | 139.3 | 167.1 | 278.5 |
All values in ps. Read across a row and the delay grows with load. Read down a column and it grows with input slew.
Looking up a value that is not in the table
Say the input slew is 80 ps and the load is 10 fF. Neither is in the table. So the tool takes the four entries around the point - in bold above - and works in two steps.
- Where is 10 fF? Between 4 and 16 fF, exactly halfway: a fraction of 0.50.
- Where is 80 ps? Between 50 and 150 ps, three-tenths of the way: a fraction of 0.30.
- Along the load, at 50 ps: 29.1 + (54.4 - 29.1) x 0.50 = 41.75 ps.
- Along the load, at 150 ps: 54.3 + (80.3 - 54.3) x 0.50 = 67.30 ps.
- Then along the slew: 41.75 + (67.30 - 41.75) x 0.30 = 49.41 ps.
That is the cell's delay: 49.41 ps. The smooth curve the table was sampled from gives 48.88 ps there. Straight lines between samples are an estimate, and the error is why libraries put their rows and columns close together where cells usually work.
The slew table, and slew passed along
The library stores the output slew the same way, in a second table with the same axes. That is how the tool passes slew down a path.
Take gate A, a NAND2_X1 driving the fanout-16 load, with a sharp 20 ps input. Its transition table gives an output slew of 198.9 ps. The 2.2 x R x C estimate from 2.3 said 191.2 ps; the table also includes the gate's own switching and its input edge.
Now gate B, another NAND2_X1 driving 8 fF, receives that edge:
| Gate B's input edge | Gate B's delay |
|---|---|
| A sharp 20 ps edge | 30.9 ps |
| The 199 ps edge from gate A | 79.7 ps |
Gate B is 48.8 ps slower, and nothing about gate B changed. The slow edge was gate A's fault.
Why it is called non-linear
The table is non-linear because the delay does not follow one straight line across the whole grid. Look down the 64 fF column: 145.5, 155.7, 184.5, 278.5. The steps grow. A single formula would miss that, so the library keeps the measurements, and the tool only draws straight lines across one small cell of the grid at a time. Newer libraries store even more detail, such as the whole shape of the current, but they are read the same way.
For how these tables sit inside a real Liberty file, and how the same idea scales to a whole standard cell library, see ASIC Volume 1.3.
Interpolating along both axes at once, by averaging the four corners. That only works at the exact centre of the cell. Always go along one axis at the two rows, then along the other.
From the table, what is the NAND2_X1 delay at an input slew of 150 ps and a load of 4 fF?
Show the answer
Answer: D. Both numbers are on the grid, so no interpolation is needed: row 150 ps, column 4 fF, 54.3 ps.
2.5 Flip-flop timing: clock-to-Q, setup and hold
A flip-flop brings three numbers to a path: clock-to-Q, setup and hold. All three are measured, not invented, and setup time turns out to be a choice about how much slowdown to accept.
Clock-to-Q is a gate delay too
When the clock edge arrives, the flip-flop's output stage drives Q like any other gate. So clock-to-Q grows with load, exactly as in 2.1. For our example flip-flop:
| Load on Q | Clock-to-Q |
|---|---|
| 2 fF | 72.8 ps |
| 8 fF | 81.1 ps |
| 32 fF | 114.4 ps |
A flip-flop driving a big fanout starts every path late. That is why heavily loaded flip-flop outputs are often buffered.
Where setup time comes from
Here is the surprising part. A flip-flop does not suddenly fail when the data is a little late. As the data arrives closer to the clock edge, the flip-flop takes longer and longer to decide. Only very close to the edge does it fail completely.
| Data arrives before the edge | Clock-to-Q | Slower by |
|---|---|---|
| 200 ps | 80.0 ps | 0% |
| 100 ps | 80.2 ps | 0% |
| 60 ps | 82.8 ps | 3% |
| 50 ps | 85.4 ps | 7% |
| 44.1 ps | 88.0 ps | 10% |
| 30 ps | 100.5 ps | 26% |
| 21 ps | 117.4 ps | 47% |
| 20 ps or less | fails to capture | - |
So where is "the" setup time? The library picks a rule. A common one is: setup time is where clock-to-Q has grown by 10%. For this flip-flop that is 44.1 ps. With a 5% rule the same flip-flop would have a setup time of 54.5 ps.
Setup time is not a cliff. It is a line drawn on a slope, at a place where the slowdown is small enough to live with. Arrive earlier than it and the flip-flop behaves exactly as the library promises.
Hold time, and why it can be negative
Hold time is measured the same way, from the other side: how soon after the edge the data may change without disturbing the captured value.
Inside the flip-flop, the clock takes a little while to reach the part that closes. So the data may sometimes be allowed to change slightly before the edge. The library then lists a negative hold time. Take a setup time of 44 ps and a hold time of -12 ps. The data must stay still from 44 ps before the edge until 12 ps before it: a window only 32 ps wide.
A negative hold time is not a mistake in the library. It means the flip-flop's own internal clock delay gives you some hold margin for free.
Treating setup and hold as fixed properties of a flip-flop. They are read from tables too, indexed by the data slew and the clock slew. A slow clock edge at the flip-flop changes its setup and hold.
A library defines setup time as the point where clock-to-Q has grown by 10%. If it switched to a 5% rule, what would happen to the setup time?
Show the answer
Answer: B. A 5% slowdown happens earlier on the curve, further from the edge, so the data has to arrive earlier. For this flip-flop the setup time grows from 44.1 ps to 54.5 ps. A stricter rule means a longer setup time.
What you learned
- A gate's delay is its own delay plus about 0.69 x R x C, so more load means more delay.
- A bigger gate charges its load faster, but loads the gate before it more.
- A wire's own delay grows with the square of its length; repeaters cut it down.
- Slew is the time an edge takes, and a slow edge makes the next gate slower too.
- Libraries store delay and slew as tables of input slew against load, read by interpolation.
- Setup time is where clock-to-Q has slowed by an agreed amount, not a sudden cliff.
- Hold time can be negative, because of the flip-flop's own internal clock delay.
Key words from this volume
Every word below has a plain-English entry in the glossary.
- Capacitance
- Resistance
- Transistor
- Fanout
- Drive strength
- PVT corner
- RC delay
- Buffer (gate)
- Repeater
- Slew (transition time)
- Timing library (Liberty file)
- Interpolation
- NLDM table
Practice
Upsize or not?
A NAND2_X1 drives a fanout of 1 (2.5 fF) and takes 15.2 ps. An engineer swaps it for a NAND2_X4, which takes 13.3 ps on the same load. The X1 driving it used to see 2.5 fF; now it sees 7.0 fF. Is the path faster?
Show the solution
Add both gates. Before: the driving X1 takes 15.2 ps into an X1 input, then the X1 takes 15.2 ps, so about 30.4 ps. After: the driving X1 takes 24.6 ps into the bigger X4 input, then the X4 takes 13.3 ps, so about 37.9 ps.
The path is slower by about 7.5 ps. The X4 saved 1.9 ps on a light load, and cost the gate before it 9.4 ps. Upsizing pays on heavy loads, not light ones.
How many repeaters?
In the 4 mm wire table, the delay falls from 1802.8 ps to 868.1 ps as the pieces go from 1 to 8, then rises again. Explain in two sentences why it rises.
Show the solution
Each extra piece shortens the wires, which saves wire delay, but it also adds one more buffer, with its own 15 ps and its own driver term. Past about eight pieces the wires are short enough that their square-law term is small, and the buffers' fixed cost wins.
Interpolate by hand
Use the NAND2_X1 table to find the delay at an input slew of 100 ps and a load of 10 fF.
Show the solution
The four corners are the same bold ones: 29.1 and 54.4 at 50 ps, 54.3 and 80.3 at 150 ps.
- 10 fF is halfway between 4 and 16 fF, so along the load: 41.75 ps at 50 ps and 67.30 ps at 150 ps.
- 100 ps is halfway between 50 and 150 ps, a fraction of 0.50.
- 41.75 + (67.30 - 41.75) x 0.50 = 54.52 ps.
Only the slew fraction changed from the worked example: 0.50 instead of 0.30.
Blame the right gate
A timing report shows a small NAND gate taking 79.7 ps, far more than you expect for its 8 fF load. What would you look at before resizing that gate?
Show the solution
Its input slew. From the table, the same gate on 8 fF takes 30.9 ps with a sharp input edge. It takes 79.7 ps when its input is a 199 ps ramp from a weak gate driving a heavy load.
So look at the gate driving it. Upsizing that one, or splitting its load, sharpens the edge and fixes the slow gate without touching it. Real timing reports print the input transition beside every delay for exactly this reason.
Interview corner
Why does wire delay scale with length squared?
"Why does the delay of a long wire grow with the square of its length, and what do you do about it?"
Show the solution
"A wire is a distributed RC line. Both its resistance and its capacitance are proportional to its length, and its own delay goes with their product, roughly 0.38 R C. Double the length and each doubles, so that term goes up four times.
The fix is repeaters: cut the wire into segments and drive each one with a buffer. Each segment has a quarter of the square-law delay, and you pay one buffer delay per segment. There is an optimum number of segments, and physical design tools insert them automatically on long nets."
What is an NLDM table?
"How does a timing tool know the delay of a standard cell?"
Show the solution
"From the Liberty file. Each timing arc of each cell has a delay table and a transition table, indexed by the input slew and the output load, measured by characterising the cell in SPICE. The tool computes the actual slew and load at that instance, finds the four surrounding entries, and interpolates - first along one axis at both neighbouring values, then along the other.
The output transition from that lookup becomes the input slew for the next cell, so slew propagates along the path along with the arrival time."
Volume 03 names the paths themselves. Where may a path start and end? What are the four kinds every tool sorts them into? And how does a timing arc say which input edge causes which output edge?