Volume 02 Beginner 5 sub-modules ~25 min read

Where Delay Comes From

So far every delay has been a number handed to you. This volume shows where the numbers come from. A gate is slow because it has to charge a capacitance through a resistance. A wire is slower still, because it is made of both. And a timing tool knows all of it from tables - which you will learn to read by hand.

You will learn
  • Why a gate's delay grows with its load, and what a bigger gate buys you
  • Why wire delay grows with the square of the length, and how repeaters fix it
  • What slew is, and why a slow edge slows down the next gate
  • How to read and interpolate a library delay table by hand
  • Where a flip-flop's clock-to-Q, setup and hold numbers really come from
You need

2.1 Gate delay and what changes it

A gate is slow because it has to charge a capacitance through a resistance. More load means more to charge, so more delay. A bigger gate has less resistance, so it charges the same load faster.

Inside every gate are transistors: tiny switches. When the input changes, one of them switches on and starts pushing current into the output wire. That wire, and every gate input hanging off it, has to fill up with charge before the voltage reaches the other side.

A switched-on transistor is not a perfect switch. It has resistance, so the current is limited. A little current filling a big capacitance takes a long time.

A gate modelled as a resistance charging a capacitance the gate R: its switch resistance C: everything it must charge to the next gates delay = the gate's own delay + 0.69 x R x C
Figure 2.1 - The whole of gate delay in one picture. The gate's switched-on transistor behaves like a resistance R. Everything it drives - the wire and the inputs of the next gates - behaves like a capacitance C. The output rises as C charges through R.
Gate delay, first-order delay = (the gate's own delay) + 0.69 x R x C

The 0.69 comes from how a capacitance charges: it reaches half the supply voltage after 0.69 of the time R x C. With R in kilo-ohms and C in femtofarads, R x C comes out in picoseconds.

One gate, three loads, two sizes

This course uses a small made-up library, with numbers chosen to keep the sums readable. Its NAND gate comes in two sizes. The bigger one, X4, has a quarter of the resistance, and four times the input capacitance.

Load Capacitance NAND2_X1 NAND2_X4
Fanout 1 2.5 fF 15.2 ps 13.3 ps
Fanout 4 8.0 fF 26.6 ps 16.2 ps
Fanout 16 29.0 fF 70.3 ps 27.1 ps

For the X1: 10 ps + 0.69 x 3.00 kOhm x 29.0 fF = 70.3 ps. For the X4: 12 ps + 0.69 x 0.75 kOhm x 29.0 fF = 27.1 ps. On a light load the two are nearly the same. On a heavy load the big gate is more than twice as fast.

Nothing is free: the big gate's input

A bigger gate has bigger transistors, and so a bigger input capacitance. The gate before it now has more to charge. That is the price of drive strength.

Two gates in a row, fanout 16 at the end First gate Second gate Total
X1 then X1 15.2 ps 70.3 ps 85.5 ps
X1 then X4 24.6 ps 27.1 ps 51.6 ps

The first gate got slower, from 15.2 ps to 24.6 ps. The path still got much faster overall. Choosing gate sizes is this trade-off, made millions of times by the synthesis tool.

What else changes a gate's delay

  1. Temperature. A hot chip is usually slower.
  2. Supply voltage. A lower voltage pushes less current, so it is slower.
  3. Manufacturing. No two chips are identical; some come out slow, some fast.
  4. The input edge. A slowly rising input makes the gate slower. Sub-module 2.3 explains why.

The first three are gathered into PVT corners. For the fanout-4 X1 in this example library, that gives 18.6 ps at the fast corner, 26.6 ps typically, and 38.6 ps at the slow one. Volume 11 is all about them.

Common mistake

Thinking a bigger gate is always faster. It drives its own load faster, but it loads the gate before it more heavily. Upsize a gate whose load is already small and the path can get slower.

Quick check

A gate has an own delay of 8 ps and a resistance of 2 kOhm. It drives 10 fF. What is its delay?

Show the answer

Answer: B. Delay = 8 + 0.69 x 2 x 10 = 8 + 13.9 = 21.9 ps. The trap answers forget the 0.69, or forget the gate's own delay.

2.2 Wire delay: resistance and capacitance

A wire has resistance and capacitance along its whole length. Make it twice as long and both double, so its own delay goes up four times. Long wires are slow in a way gates never are.

On a small chip, wires are cheap. On a big one, a signal may cross several millimetres, and the wire costs more time than the gates at either end.

The wire is not one resistance and one capacitance. It is many tiny pieces of each, spread out along its length. The charge has to travel through the resistance of the near end to fill the capacitance of the far end.

A driver, a wire and a load delay = 0.69 R_d (C_w + C_L) + 0.38 R_w C_w + 0.69 R_w C_L

Here R_d is the driver, R_w and C_w are the wire, and C_L is the load.

Read it term by term. The driver charges everything. The wire charges itself, and 0.38 is the figure for charge spread along a line. Finally, the wire's resistance slows the charging of the load at the far end.

A wire, measured

Take a wire of 1 ohm and 0.2 fF per micrometre, driven through 1 kOhm into a 5 fF load.

Length R wire C wire Driver term Wire term Load term Total
100 um 0.10 kOhm 20 fF 17.3 0.8 0.3 18.4 ps
500 um 0.50 kOhm 100 fF 72.8 19.0 1.7 93.5 ps
1000 um 1.00 kOhm 200 fF 142.1 76.0 3.5 221.6 ps
2000 um 2.00 kOhm 400 fF 280.7 304.0 6.9 591.7 ps
4000 um 4.00 kOhm 800 fF 558.0 1216.0 13.9 1787.8 ps

Watch the wire term: 76.0, then 304.0, then 1216.0. Each doubling of the length multiplies it by four. At 4 mm it is two-thirds of the total.

Delay of a driven wire against its length, with the wire's own term shown separately 0 500 1000 1500 2000 2500 3000 3500 4000 0 400 800 1200 1600 2000 wire length (micrometres) delay (ps) total delay wire term alone
Figure 2.2 - The total delay (gold) bends upwards because the wire's own term (purple) grows with the square of the length. Up to half a millimetre the driver dominates; by two millimetres the wire itself does.
In plain words

RC delay grows as length times length. A wire twice as long is not twice as slow - it is up to four times as slow. Nothing else on a chip behaves like that.

The fix: cut the wire into pieces

If a wire's own delay grows with the square of its length, then several short wires are faster than one long one. So long wires are cut into pieces, with a buffer driving each piece. Those buffers are called repeaters.

The same 4 mm wire, cut into equal pieces, with a buffer of 15 ps, 1.0 kOhm and 5 fF before each:

Pieces Each piece Total delay
1 4000 um 1802.8 ps
2 2000 um 1213.3 ps
4 1000 um 946.2 ps
8 500 um 868.1 ps
12 333 um 891.3 ps
16 250 um 939.8 ps
32 125 um 1197.3 ps

Eight pieces bring 1802.8 ps down to 868.1 ps, less than half. Past that, the buffers' own delay starts to win, and more pieces make it worse again. There is always a best number.

Common mistake

Adding repeaters to every wire. On a short wire the wire term is tiny - 0.8 ps at 100 um - so a repeater only adds its own delay. Repeaters pay only on wires long enough for the square law to bite.

Quick check

A wire's own term is 19.0 ps at 500 um. Nothing else changes. What is it at 1500 um?

Show the answer

Answer: C. The wire term grows with the square of the length. Three times the length gives nine times the term: 9 x 19.0 = 171.0 ps. The first answer grows it in a straight line, which is exactly the mistake the square law punishes.

2.3 Slew, the transition time

Slew is how long an edge takes to get from low to high. A gate with a heavy load makes a slow edge, and a slow edge arriving at the next gate makes that gate slow too. Delay is passed down the path.

So far every edge has been drawn as a straight vertical line. Real edges are ramps. When a gate charges a big capacitance, its output climbs slowly.

A sharp edge and a slow edge, each with its 10% and 90% points marked light load: a sharp edge heavy load: a slow edge 90% 10% a short slew a long slew
Figure 2.3 - The same swing, two speeds. Slew is measured between the 10% and 90% points, because the very start and end of an edge creep along and are hard to pin down.

The slew of an edge charging through R is about 2.2 x R x C, measured from 10% to 90%.

Gate Load C Output slew
NAND2_X1 fanout 1 2.5 fF 16.5 ps
NAND2_X1 fanout 4 8.0 fF 52.7 ps
NAND2_X1 fanout 16 29.0 fF 191.2 ps
NAND2_X4 fanout 16 29.0 fF 47.8 ps

Why a slow edge slows the next gate

The next gate does not switch the instant its input starts to move. Its transistors only turn fully on once the input is well past halfway. A slow ramp spends a long time getting there, so the next gate starts late - and, part-way on, it switches weakly.

That means delay is not just a property of each gate on its own. A weak gate driving a big load hurts itself and the gate after it. Sub-module 2.4 puts a number on this.

Remember

A timing tool carries two numbers along every path: the arrival time, and the slew of that arrival. Each gate's delay depends on the slew coming in, and each gate produces a new slew going out.

Common mistake

Fixing a slow gate by looking only at that gate. If its input edge is slow, the real cure is often upstream: a stronger gate before it, or less load on the net feeding it.

Quick check

A gate with 1.5 kOhm of resistance drives 20 fF. Roughly what is its output slew?

Show the answer

Answer: A. Slew is about 2.2 x R x C = 2.2 x 1.5 x 20 = 65.9 ps, so about 66 ps. The 21 ps answer uses 0.69 instead of 2.2 - that is the delay factor, not the slew factor.

2.4 Library delay tables made simple

A timing library stores each cell's delay as a small table, looked up by input slew and output load. Between the table's rows and columns, the tool interpolates in straight lines.

The formula in 2.1 gives the right shape, but real transistors do not follow it exactly. So the people who make a cell library measure every cell, at many slews and many loads, and store the results.

The result is an NLDM table: a grid of numbers. Here is our NAND2_X1's delay.

Input slew \ load 1 fF 4 fF 16 fF 64 fF
10 ps 14.1 20.4 45.4 145.5
50 ps 22.7 29.1 54.4 155.7
150 ps 47.8 54.3 80.3 184.5
400 ps 132.3 139.3 167.1 278.5

All values in ps. Read across a row and the delay grows with load. Read down a column and it grows with input slew.

Looking up a value that is not in the table

Say the input slew is 80 ps and the load is 10 fF. Neither is in the table. So the tool takes the four entries around the point - in bold above - and works in two steps.

  1. Where is 10 fF? Between 4 and 16 fF, exactly halfway: a fraction of 0.50.
  2. Where is 80 ps? Between 50 and 150 ps, three-tenths of the way: a fraction of 0.30.
  3. Along the load, at 50 ps: 29.1 + (54.4 - 29.1) x 0.50 = 41.75 ps.
  4. Along the load, at 150 ps: 54.3 + (80.3 - 54.3) x 0.50 = 67.30 ps.
  5. Then along the slew: 41.75 + (67.30 - 41.75) x 0.30 = 49.41 ps.

That is the cell's delay: 49.41 ps. The smooth curve the table was sampled from gives 48.88 ps there. Straight lines between samples are an estimate, and the error is why libraries put their rows and columns close together where cells usually work.

The slew table, and slew passed along

The library stores the output slew the same way, in a second table with the same axes. That is how the tool passes slew down a path.

Take gate A, a NAND2_X1 driving the fanout-16 load, with a sharp 20 ps input. Its transition table gives an output slew of 198.9 ps. The 2.2 x R x C estimate from 2.3 said 191.2 ps; the table also includes the gate's own switching and its input edge.

Now gate B, another NAND2_X1 driving 8 fF, receives that edge:

Gate B's input edge Gate B's delay
A sharp 20 ps edge 30.9 ps
The 199 ps edge from gate A 79.7 ps

Gate B is 48.8 ps slower, and nothing about gate B changed. The slow edge was gate A's fault.

Why it is called non-linear

The table is non-linear because the delay does not follow one straight line across the whole grid. Look down the 64 fF column: 145.5, 155.7, 184.5, 278.5. The steps grow. A single formula would miss that, so the library keeps the measurements, and the tool only draws straight lines across one small cell of the grid at a time. Newer libraries store even more detail, such as the whole shape of the current, but they are read the same way.

For how these tables sit inside a real Liberty file, and how the same idea scales to a whole standard cell library, see ASIC Volume 1.3.

Common mistake

Interpolating along both axes at once, by averaging the four corners. That only works at the exact centre of the cell. Always go along one axis at the two rows, then along the other.

Quick check

From the table, what is the NAND2_X1 delay at an input slew of 150 ps and a load of 4 fF?

Show the answer

Answer: D. Both numbers are on the grid, so no interpolation is needed: row 150 ps, column 4 fF, 54.3 ps.

2.5 Flip-flop timing: clock-to-Q, setup and hold

A flip-flop brings three numbers to a path: clock-to-Q, setup and hold. All three are measured, not invented, and setup time turns out to be a choice about how much slowdown to accept.

Clock-to-Q is a gate delay too

When the clock edge arrives, the flip-flop's output stage drives Q like any other gate. So clock-to-Q grows with load, exactly as in 2.1. For our example flip-flop:

Load on Q Clock-to-Q
2 fF 72.8 ps
8 fF 81.1 ps
32 fF 114.4 ps

A flip-flop driving a big fanout starts every path late. That is why heavily loaded flip-flop outputs are often buffered.

Where setup time comes from

Here is the surprising part. A flip-flop does not suddenly fail when the data is a little late. As the data arrives closer to the clock edge, the flip-flop takes longer and longer to decide. Only very close to the edge does it fail completely.

Data arrives before the edge Clock-to-Q Slower by
200 ps 80.0 ps 0%
100 ps 80.2 ps 0%
60 ps 82.8 ps 3%
50 ps 85.4 ps 7%
44.1 ps 88.0 ps 10%
30 ps 100.5 ps 26%
21 ps 117.4 ps 47%
20 ps or less fails to capture -
Clock-to-Q against how long before the clock edge the data arrived 0 20 40 60 80 100 120 70 80 90 100 110 120 130 data arrives this long before the clock edge (ps) clock-to-Q (ps) 80 ps plus 10% setup time fails
Figure 2.4 - Far from the edge, clock-to-Q is a steady 80 ps. As the data arrives later, the flip-flop takes longer to settle, and at 20 ps it cannot decide at all. The library calls the point where clock-to-Q has grown by 10% the setup time: here 44.1 ps.

So where is "the" setup time? The library picks a rule. A common one is: setup time is where clock-to-Q has grown by 10%. For this flip-flop that is 44.1 ps. With a 5% rule the same flip-flop would have a setup time of 54.5 ps.

In plain words

Setup time is not a cliff. It is a line drawn on a slope, at a place where the slowdown is small enough to live with. Arrive earlier than it and the flip-flop behaves exactly as the library promises.

Hold time, and why it can be negative

Hold time is measured the same way, from the other side: how soon after the edge the data may change without disturbing the captured value.

Inside the flip-flop, the clock takes a little while to reach the part that closes. So the data may sometimes be allowed to change slightly before the edge. The library then lists a negative hold time. Take a setup time of 44 ps and a hold time of -12 ps. The data must stay still from 44 ps before the edge until 12 ps before it: a window only 32 ps wide.

Remember

A negative hold time is not a mistake in the library. It means the flip-flop's own internal clock delay gives you some hold margin for free.

Common mistake

Treating setup and hold as fixed properties of a flip-flop. They are read from tables too, indexed by the data slew and the clock slew. A slow clock edge at the flip-flop changes its setup and hold.

Quick check

A library defines setup time as the point where clock-to-Q has grown by 10%. If it switched to a 5% rule, what would happen to the setup time?

Show the answer

Answer: B. A 5% slowdown happens earlier on the curve, further from the edge, so the data has to arrive earlier. For this flip-flop the setup time grows from 44.1 ps to 54.5 ps. A stricter rule means a longer setup time.

What you learned

Key words from this volume

Every word below has a plain-English entry in the glossary.

Practice

Practice 1

Upsize or not?

A NAND2_X1 drives a fanout of 1 (2.5 fF) and takes 15.2 ps. An engineer swaps it for a NAND2_X4, which takes 13.3 ps on the same load. The X1 driving it used to see 2.5 fF; now it sees 7.0 fF. Is the path faster?

Show the solution

Add both gates. Before: the driving X1 takes 15.2 ps into an X1 input, then the X1 takes 15.2 ps, so about 30.4 ps. After: the driving X1 takes 24.6 ps into the bigger X4 input, then the X4 takes 13.3 ps, so about 37.9 ps.

The path is slower by about 7.5 ps. The X4 saved 1.9 ps on a light load, and cost the gate before it 9.4 ps. Upsizing pays on heavy loads, not light ones.

Practice 2

How many repeaters?

In the 4 mm wire table, the delay falls from 1802.8 ps to 868.1 ps as the pieces go from 1 to 8, then rises again. Explain in two sentences why it rises.

Show the solution

Each extra piece shortens the wires, which saves wire delay, but it also adds one more buffer, with its own 15 ps and its own driver term. Past about eight pieces the wires are short enough that their square-law term is small, and the buffers' fixed cost wins.

Practice 3

Interpolate by hand

Use the NAND2_X1 table to find the delay at an input slew of 100 ps and a load of 10 fF.

Show the solution

The four corners are the same bold ones: 29.1 and 54.4 at 50 ps, 54.3 and 80.3 at 150 ps.

  1. 10 fF is halfway between 4 and 16 fF, so along the load: 41.75 ps at 50 ps and 67.30 ps at 150 ps.
  2. 100 ps is halfway between 50 and 150 ps, a fraction of 0.50.
  3. 41.75 + (67.30 - 41.75) x 0.50 = 54.52 ps.

Only the slew fraction changed from the worked example: 0.50 instead of 0.30.

Practice 4

Blame the right gate

A timing report shows a small NAND gate taking 79.7 ps, far more than you expect for its 8 fF load. What would you look at before resizing that gate?

Show the solution

Its input slew. From the table, the same gate on 8 fF takes 30.9 ps with a sharp input edge. It takes 79.7 ps when its input is a 199 ps ramp from a weak gate driving a heavy load.

So look at the gate driving it. Upsizing that one, or splitting its load, sharpens the edge and fixes the slow gate without touching it. Real timing reports print the input transition beside every delay for exactly this reason.

Interview corner

Interview question 1

Why does wire delay scale with length squared?

"Why does the delay of a long wire grow with the square of its length, and what do you do about it?"

Show the solution

"A wire is a distributed RC line. Both its resistance and its capacitance are proportional to its length, and its own delay goes with their product, roughly 0.38 R C. Double the length and each doubles, so that term goes up four times.

The fix is repeaters: cut the wire into segments and drive each one with a buffer. Each segment has a quarter of the square-law delay, and you pay one buffer delay per segment. There is an optimum number of segments, and physical design tools insert them automatically on long nets."

Interview question 2

What is an NLDM table?

"How does a timing tool know the delay of a standard cell?"

Show the solution

"From the Liberty file. Each timing arc of each cell has a delay table and a transition table, indexed by the input slew and the output load, measured by characterising the cell in SPICE. The tool computes the actual slew and load at that instance, finds the four surrounding entries, and interpolates - first along one axis at both neighbouring values, then along the other.

The output transition from that lookup becomes the input slew for the next cell, so slew propagates along the path along with the arrival time."

Volume 03 names the paths themselves. Where may a path start and end? What are the four kinds every tool sorts them into? And how does a timing arc say which input edge causes which output edge?