Volume 03 Beginner 5 sub-modules ~20 min read

Timing Paths

A timing tool never looks at a design as a whole. It breaks it into paths - millions of them - each with one start, one end and a chain of delays between. This volume shows where paths may start and end, the four kinds every report sorts them into, how rising and falling edges are followed through each gate, and why the clock gets a path of its own.

You will learn
  • Where a timing path can start and where it can end
  • The four path types, and how each one is timed
  • Which clock edge launches the data and which one captures it
  • What a timing arc is, and why rise and fall are followed separately
  • Why the clock has its own path, and why only the difference in clock arrival matters
You need

3.1 Start points and end points

Every timing path begins where data is launched and ends where data is captured. There are only two kinds of beginning and two kinds of end, and that is what lets a tool find every path in a chip.

A timing tool does not try to understand your design. It breaks it into pieces it can time one at a time: paths. Each path has exactly one start point and one endpoint.

A path may start at A path may end at
The clock pin of a flip-flop, CK The data pin of a flip-flop, D
An input port of the design An output port of the design

Why the clock pin and not the Q output? Because the data does not leave the flip-flop until the clock edge arrives there. The path really begins with the clock edge, and the clock-to-Q delay is its first step.

Every path in a small design

Here is a design small enough to list completely: two inputs, three flip-flops and one output. U1 is an AND gate, U2 an XOR gate, U3 an OR gate, and FF3's output also feeds FF1 directly, with no logic at all.

# Path Type
1 in_a -> U1 -> FF2/D input to register
2 in_b -> U2 -> FF3/D input to register
3 in_b -> U3 -> out_y input to output
4 FF1/CK -> U1 -> FF2/D register to register
5 FF2/CK -> U2 -> FF3/D register to register
6 FF3/CK -> U3 -> out_y register to output
7 FF3/CK -> FF1/D register to register

Seven paths, from five start points (FF1/CK, FF2/CK, FF3/CK, in_a, in_b) to four endpoints (FF1/D, FF2/D, FF3/D, out_y). The tool times each one separately, and each gets its own slack.

Notice path 7. It has no logic at all, just a wire. It still counts, and it is exactly the kind of path that fails hold, because nothing slows it down.

In plain words

Start points launch, endpoints capture. The tool walks forward from every start point, follows every branch, and stops each time it reaches an endpoint. Every walk is one path.

Remember

A real chip has millions of start points and millions of endpoints, and the number of paths between them is far larger. The tool does not list them all. It works out the worst arrival at each endpoint cleverly, then reports the worst paths.

Quick check

Which of these can be the start point of a timing path?

Show the answer

Answer: C. Paths start where data is launched: at a flip-flop's clock pin, or at an input port. The data leaves through Q, but only because the clock edge reached CK first, so the path starts at CK.

3.2 The four path types

Two kinds of start and two kinds of end make four kinds of path. Every report sorts its paths into these four, and each is timed with the same arithmetic.

The four path types in one design: input to register, register to register, register to output, and input to output the design in in out out logic logic logic logic FF1 FF2 clk 1 input to register 2 register to register 3 register to output 4 input to output
Figure 3.1 - The four kinds of path. A path starts at an input port or a flip-flop's clock, and ends at a flip-flop's data pin or an output port. Every path in every chip is one of these four.

All four use the same sum as Volume 01: the data arrives at some time, it is needed by some time, and the difference is the slack. What changes is where the two times come from.

Type Where the arrival starts Where the requirement comes from
Register to register The launch flip-flop's clock-to-Q The capture flip-flop's setup time
Input to register The input delay outside the design The capture flip-flop's setup time
Register to output The launch flip-flop's clock-to-Q The output delay outside the design
Input to output The input delay The output delay

An input delay says: the outside world launched this data on the same clock, and it has already spent this long getting here. An output delay says: the outside world needs this much time after our output to reach its own flip-flop. Volume 10 covers both properly.

All four, timed at 100 MHz

Type Arrival Required Slack
Register to register 4.55 ns 9.85 ns 5.30 ns
Input to register 3.44 ns 9.85 ns 6.41 ns
Register to output 1.41 ns 6.00 ns 4.59 ns
Input to output 3.06 ns 6.00 ns 2.94 ns
  1. Input to register. Input delay 3.00 + AND2 0.14 + wire 0.30 = 3.44 ns. Needed by 10 - 0.15 = 9.85 ns.
  2. Register to output. Clock-to-Q 0.35 + OR2 0.16 + output pad 0.90 = 1.41 ns. Needed by 10 - 4.00 = 6.00 ns.
  3. Input to output. Input delay 2.00 + OR2 0.16 + output pad 0.90 = 3.06 ns. Needed by 6.00 ns.
Common mistake

Forgetting that input and output paths are only half inside your design. The tool cannot see the other chip; it only knows the input and output delays you gave it. Leave them out and those paths are not checked at all - and the report will not complain.

Quick check

The clock is 10 ns. An input arrives 6.50 ns after the edge, then passes 2.80 ns of logic to a flip-flop with a 0.20 ns setup time. What is the slack?

Show the answer

Answer: A. The data arrives at 6.50 + 2.80 = 9.30 ns and is needed by 10 - 0.20 = 9.80 ns. Slack = 9.80 - 9.30 = 0.50 ns. An input path is timed exactly like any other, except that its first step is the input delay rather than a clock-to-Q.

3.3 Launch edge and capture edge

Every check compares two clock edges: the one that launched the data and the one that captures it. For an ordinary path they are one period apart for setup, and the very same edge for hold.

The launch edge sends the data. The capture edge is the one the setup check races against. Choosing the right pair of edges is the first thing a tool does for every path.

The launch edge and the capture edge of an ordinary path, one period apart 0 1 2 3 clk ff1_q ff2_q launch edge capture edge
Figure 3.2 - An ordinary path on one clock. Edge 1 launches the data from FF1; edge 2, one period later, captures it at FF2. The hold check uses edge 1 at both ends: the new data must not reach FF2 while edge 1 is still being captured there.

Three ways to pair the edges on one clock

A flip-flop can act on the rising edge, or on the falling one. That gives different pairs:

Launch on Capture on Setup: time between the edges Hold: edge pair
Rising Rising 10.00 ns (0 to 10) 0 and 0: the same edge
Rising Falling 5.00 ns (0 to 5) 0 and -5: half a period earlier
Falling Rising 5.00 ns (5 to 10) 5 and 0: half a period earlier

Clock of 10 ns with a 50% duty cycle. A path from a rising-edge flop to a falling-edge flop gets only half a period for setup. But its hold check becomes easy, because the capture edge it is checked against came half a period before the launch.

In plain words

For setup, the tool looks for the first capture edge after the launch edge. For hold, it looks one capture edge earlier than that. Everything in Volume 07 - divided clocks, odd ratios, inverted clocks - is this same rule applied to trickier waveforms.

Common mistake

Assuming every path gets a full clock period. A path between a rising-edge and a falling-edge flop gets half. Designs that mix edges have to be read with care, and Volume 07 works through every case.

Quick check

A path launches on a rising edge and captures on the falling edge of the same 8 ns clock (50% duty). How long does setup get?

Show the answer

Answer: B. The first falling edge after a rising edge at 0 comes at half the period, 4 ns. So the setup check compares edges 4 ns apart.

3.4 Timing arcs and unateness

A timing arc says which input change causes which output change, and how long it takes. Rising and falling edges have different delays, so a tool follows both, all the way along the path.

Inside a gate, pulling the output up and pulling it down are done by different transistors. They are not equally strong, so a gate's rising delay and falling delay differ. The library stores both.

And not every input edge causes the same output edge. That relationship has a name, unateness.

Unateness Examples A rising input causes
Positive unate Buffer, AND, OR A rising output
Negative unate Inverter, NAND, NOR A falling output
Non-unate XOR, XNOR, multiplexer select Either, depending on the other inputs
An inverter: its output falls when its input rises, a short delay later a y
Figure 3.3 - An inverter is negative unate. Each rising input edge produces a falling output edge, and each falling input edge a rising one, a short delay later. There is no clock here - this is pure combinational delay.

Following rise and fall down a chain

Take a path through four gates, starting from a flip-flop whose Q rises at 0.36 ns or falls at 0.33 ns after the clock edge. Each gate has its own rise and fall delay.

After Rise arrival Fall arrival How
Launch flop Q 0.36 ns 0.33 ns clock-to-Q
INV (rise 0.05, fall 0.03) 0.38 ns 0.39 ns flips: rise comes from the fall
NAND2 (rise 0.08, fall 0.06) 0.47 ns 0.44 ns flips again
XOR2 (rise 0.12, fall 0.11) 0.59 ns 0.58 ns either input edge, so take the later
BUF (rise 0.07, fall 0.08) 0.66 ns 0.66 ns keeps the direction

The latest arrival at the end is 0.66 ns. For hold, the tool runs the same walk taking the earlier edge at the XOR, and gets 0.63 ns.

Why not just add the worst delay of each gate?

Adding each gate's slower delay, ignoring which edge is which, gives 0.36 + 0.05 + 0.08 + 0.12 + 0.08 = 0.69 ns. That is 0.03 ns worse than any real edge can be. An edge cannot be a rise and a fall at the same time, and following the arcs knows that.

Remember

A tool always follows the arcs. Pessimism of a few picoseconds per path sounds small. Summed over millions of paths, it is the difference between closing timing and adding a week of work.

Common mistake

Assuming a gate has one delay. Every arc has a rise delay and a fall delay, and each of those is a table indexed by slew and load, as in Volume 02.

Quick check

A flip-flop's Q rises and falls at 0.30 ns. It drives a NAND2 (rise 0.10, fall 0.07), then a NOR2 (rise 0.12, fall 0.05). When does a rising edge arrive at the NOR2's output?

Show the answer

Answer: D. Both gates are negative unate. A rising NOR2 output comes from a falling NAND2 output, which comes from a rising Q. So: 0.30 + 0.07 = 0.37 ns, then 0.37 + 0.12 = 0.49 ns. The falling edge arrives at 0.45 ns.

3.5 Clock paths versus data paths

The clock has a path too: from its source, through the clock tree, to every flip-flop. A tool times it separately, and only the difference between the launch and capture clock arrivals changes the slack.

So far the clock edge has reached both flip-flops at the same instant. It never does. The clock starts at one place and is carried to every flip-flop by a tree of buffers and wires. That takes time: the clock latency.

A timing path with its clock tree: the clock reaches FF1 after 0.85 ns and FF2 after 0.90 ns FF1 D Q t_cq 0.35 AND2 0.14 ns 16-bit adder 3.42 ns MUX 0.64 ns setup 0.15 FF2 D Q clk period 10.00 ns 0.85 ns 0.90 ns
Figure 3.4 - The same path, now with its clock network. The launch edge reaches FF1 0.85 ns after leaving the clock source, and the capture edge reaches FF2 0.90 ns after. The data path and the clock paths are timed separately, then compared.

The arithmetic, with the clock paths in

  1. Arrival. Launch edge at 0, plus 0.85 to reach FF1, plus clock-to-Q 0.35, plus logic 4.20 = 5.40 ns.
  2. Required. Capture edge at 10, plus 0.90 to reach FF2, less setup 0.15 = 10.75 ns.
  3. Slack. 10.75 - 5.40 = 5.35 ns.

Without the clock tree the slack was 5.30 ns. With it, it is 5.35 ns: 0.05 ns better. That 0.05 ns is exactly how much later the clock reaches FF2 than FF1.

Only the difference matters

Clock reaches FF1 Clock reaches FF2 Difference Setup slack Hold slack
(ideal clock) (ideal clock) 0 5.30 ns 1.37 ns
0.85 ns 0.90 ns +0.05 5.35 ns 1.32 ns
2.85 ns 2.90 ns +0.05 5.35 ns 1.32 ns
0.90 ns 0.85 ns -0.05 5.25 ns 1.42 ns

Add two whole nanoseconds to both clock paths and nothing changes. Swap them, and setup gets worse while hold gets better. That difference is clock skew, and Volume 06 treats it in full.

In plain words

The clock is a runner too. If the capture clock arrives late, the data gets extra time for setup, but the new data also looks earlier to the hold check. Skew always helps one check and hurts the other.

Common mistake

Adding the clock latency to the data path only. The launch clock latency delays the data, and the capture clock latency delays the deadline. Both go in, on opposite sides, which is why the common part cancels.

Quick check

The clock reaches the launch flop at 1.20 ns and the capture flop at 1.10 ns. Compared with an ideal clock, what happens to setup slack?

Show the answer

Answer: C. The capture clock arrives 0.10 ns earlier than the launch clock, so the deadline moves 0.10 ns earlier relative to the data. Setup slack drops by 0.10 ns. Only the difference matters, not the 1.20 or 1.10 themselves.

What you learned

Key words from this volume

Every word below has a plain-English entry in the glossary.

Practice

Practice 1

Classify and time

A path leaves a flip-flop with a clock-to-Q of 0.40 ns and passes 1.20 ns of logic. It leaves the chip through an output port with an output delay of 7.00 ns. The clock is 10 ns. What type of path is it, and what is its slack?

Show the solution

It is a register-to-output path: it starts at a flip-flop's clock pin and ends at an output port.

The data arrives at 0.40 + 1.20 = 1.60 ns. The outside world needs 7.00 ns of the cycle, so the data must leave by 10 - 7.00 = 3.00 ns. Slack = 3.00 - 1.60 = 1.40 ns.

Practice 2

Count the paths

In the small design of sub-module 3.1, suppose U2 gains a third input from in_a. How many paths are there now, and what type is the new one?

Show the solution

One new path: in_a -> U2 -> FF3/D, which is input to register. That makes eight paths.

The rule to count by: each start point, followed along every branch, to every endpoint it reaches. Now in_a reaches two endpoints, FF2/D through U1 and FF3/D through U2.

Practice 3

Which edge arrives last?

Using the four-gate chain of sub-module 3.4, suppose the XOR2 were replaced by an AND2 (positive unate) with rise 0.12 and fall 0.11. Would the latest arrival at the end still be 0.66 ns?

Show the solution

No. With an AND2 the edges keep their direction instead of taking the later of the two.

After the NAND2, rise is 0.47 ns and fall is 0.44 ns. Through the AND2, rise becomes 0.47 + 0.12 = 0.59 ns and fall 0.44 + 0.11 = 0.55 ns. Through the buffer, rise is 0.59 + 0.07 = 0.66 ns and fall 0.55 + 0.08 = 0.63 ns.

So the latest arrival is still 0.66 ns - but now only the rising edge gets there that late. The falling edge, which the XOR had pushed to 0.66 ns, arrives at 0.63 ns.

Practice 4

Skew that helps, skew that hurts

Two flip-flops share a clock. An engineer delays the clock to the capture flip-flop by 0.20 ns to rescue a setup violation of -0.12 ns. The path's hold slack was +0.15 ns. What are the two slacks now?

Show the solution

Delaying the capture clock gives setup 0.20 ns more time: -0.12 + 0.20 = +0.08 ns. It passes.

The same delay makes the new data look 0.20 ns earlier to the hold check: 0.15 - 0.20 = -0.05 ns. Hold now fails, and a hold failure cannot be fixed by slowing the clock. The rescue traded a setup problem for a worse one.

Interview corner

Interview question 1

Where does a timing path start?

"Why does a register-to-register timing path start at the clock pin of the launch flop, and not at its Q output?"

Show the solution

"Because the data only leaves when the clock edge reaches the flop. The first delay on the path is the clock-to-Q arc, from CK to Q, and it depends on the clock's slew and on Q's load. Starting at CK lets the tool include that arc, and connect the data path to the clock path that feeds it. That is how launch clock latency gets into the arrival time."

Interview question 2

Positive, negative and non-unate

"What is unateness, and why does a timing tool care?"

Show the solution

"Unateness is the relationship between input and output edge direction on a timing arc. Buffers and AND gates are positive unate, inverters and NANDs are negative unate, and XORs are non-unate - the output direction depends on the other input.

The tool cares because rise and fall delays differ. It propagates a rise arrival and a fall arrival separately, using unateness to know which input edge feeds which output edge. That is more accurate than taking each gate's worst delay, which would combine a rising and a falling edge that can never happen together."

Volume 04 takes the register-to-register path and does setup analysis properly. Arrival time, required time and slack are built one term at a time, with ten worked problems and a real report read line by line.