Volume 01 Beginner 5 sub-modules ~25 min read

Why Chips Need Timing Analysis

Every clock edge starts a race. The data sets off through the logic, and the next clock edge sets off down the clock tree. If the data loses, the chip captures the wrong value. This volume shows both ways that race can go wrong - too slow, and too fast - with real numbers for each, and explains why no amount of simulation would find them.

You will learn
  • What is really racing on a timing path, and what the finish line is
  • What setup time means, and the deadline it puts on every path
  • What hold time means, and why slowing the clock cannot fix it
  • What a setup failure and a hold failure each cost in real silicon
  • Why static timing analysis finds these bugs and simulation does not
You need
  • Volume 00, for clocks, flip-flops, delay and units

1.1 The race between data and clock

Every clock edge starts a race. The data runs through the logic towards the next flip-flop, and the clock runs towards that same flip-flop to fetch it. Timing analysis is the story of who arrives first.

Volume 00 built the pieces: a clock that ticks, flip-flops that copy on the tick, and logic that takes time. Put them together and you get a timing path.

A path always has the same three parts. A flip-flop launches the data, some logic works on it, and another flip-flop captures it. The edge that starts it is the launch edge. The edge that ends it is the capture edge.

Think of it like this

Two runners leave at the same starting gun. One carries the data through the logic. The other is the next clock edge, running down the clock wires to the capture flip-flop. The data runner has to be there, standing still, before the clock runner arrives.

The path from Volume 00, timed

A timing path with its delays: clock-to-Q, three logic stages, and the setup time at the far end FF1 D Q t_cq 0.35 AND2 0.14 ns 16-bit adder 3.42 ns MUX 0.64 ns setup 0.15 FF2 D Q clk period 10.00 ns data ready at 4.55 ns
Figure 1.1 - The data leaves FF1 0.35 ns after the launch edge and spends 4.20 ns in the logic. So it is ready 4.55 ns after the edge, and the capture edge comes at 10.00 ns. There is time to spare.

Now do the arithmetic that this whole course is made of:

  1. The data is ready 4.55 ns after the launch edge.
  2. FF2 wants it steady 0.15 ns early, so the deadline is 10.00 - 0.15 = 9.85 ns.
  3. The difference, 9.85 - 4.55 = 5.30 ns, is the spare time.

That spare time is called slack. It is the single most important number in this course, and the rule is simple: positive slack passes, negative slack fails.

Slack slack = (the time the data was allowed) - (the time the data took)

The same path, with a bigger adder

Swap the 16-bit adder for a 32-bit one, and the logic takes 10.18 ns instead of 4.20 ns.

The same path with a 32-bit adder, where the data arrives after its deadline FF1 D Q t_cq 0.35 AND2 0.14 ns 32-bit adder 9.40 ns MUX 0.64 ns setup 0.15 FF2 D Q clk period 10.00 ns data ready at 10.53 ns - too late
Figure 1.2 - Nothing else changed: same clock, same flip-flops, same clock-to-Q. The adder alone pushes the arrival to 10.53 ns, past the 9.85 ns deadline, so the slack is negative and the path fails.

The data is ready at 10.53 ns, and the deadline was 9.85 ns. The slack is 9.85 - 10.53 = -0.68 ns. The path fails by 680 picoseconds.

In plain words

A negative slack does not mean the chip is a bit slow. It means that at this clock speed, this path captures the wrong value. The design is wrong until the number goes positive.

What the failure looks like

Data arriving in time to be captured at the next clock edge 0 1 2 3 clk q1 d2 OLD NEW q2 launch edge capture edge
Figure 1.3 - The good path. FF1 launches at edge 1, the new value reaches FF2's input partway through the cycle, and FF2 captures it at edge 2. Its output q2 changes one cycle after the launch.
Data arriving after the capture edge, so the old value is captured instead 0 1 2 3 clk q1 d2 OLD NEW q2 capture edge - data still OLD
Figure 1.4 - The failing path. The new value turns up just after edge 2, so what FF2 actually captured at edge 2 was the old one. The design carries on happily with the wrong number.
Common mistake

Expecting the chip to warn you. It cannot. A flip-flop that captures a stale value has no way of knowing, and nothing downstream can tell either. The result is simply a wrong number, produced at full speed, on some chips, at some temperatures.

Quick check

A path's data is ready 6.10 ns after the launch edge. The deadline is 7.85 ns. What is the slack?

Show the answer

Answer: B. Slack is the time allowed minus the time taken: 7.85 - 6.10 = 1.75 ns. It is positive, so the path passes with 1.75 ns to spare.

1.2 Setup time: arrive early enough

Setup time is how long the data must already be steady when the clock edge arrives. It moves the deadline earlier, and it is the reason a chip has a top speed.

A flip-flop is not a camera shutter. Inside it, the arriving value has to settle before the edge can store it reliably. If it is still moving, the flip-flop may store either value, or hang between the two for a while.

So the finish line is not the capture edge. It is the capture edge minus the setup time.

The setup deadline, for a path on one clock data must arrive by: (clock period) - (setup time)

For our path, the period is 10.00 ns and the setup time is 0.15 ns, so everything must be in place by 9.85 ns. Volume 04 adds the other terms - skew, uncertainty, and the exact place each one goes.

Speed up the clock and watch the slack go

The path does not change. The time it is allowed does.

Period Frequency Setup slack Verdict
12.00 ns 83.3 MHz 7.30 ns passes
10.00 ns 100.0 MHz 5.30 ns passes
8.00 ns 125.0 MHz 3.30 ns passes
6.00 ns 166.7 MHz 1.30 ns passes
5.00 ns 200.0 MHz 0.30 ns passes
4.70 ns 212.8 MHz 0.00 ns passes, with nothing to spare

Every row is the same sum with a different period. Shorten the period by 2 ns and the slack drops by 2 ns, exactly.

Remember

The last row is the path's own speed limit: 212.8 MHz. Below that it works, above it the slack goes negative. The slowest path in the whole chip sets the speed of the chip, and that path has a name - the critical path.

Why the answer is never "just clock it slower"

Slowing the clock does buy setup slack, and sometimes that is the right answer. But the customer wanted 200 MHz, and a chip that only reaches 93.6 MHz is a different product at a different price. Volume 14 is entirely about closing that gap.

Common mistake

Thinking the data has the whole clock period. It has the period, less the setup time at the far end, less the clock-to-Q at the near end, and less a few more terms this course adds later. On a fast design those "small" terms can eat a fifth of the cycle.

Quick check

A path arrives at 3.90 ns. The flip-flop's setup time is 0.10 ns. What is the shortest clock period it can work at?

Show the answer

Answer: C. The deadline is the period minus the setup time, and it must be at least 3.90 ns. So the period must be at least 3.90 + 0.10 = 4.00 ns, which is 250 MHz.

1.3 Hold time: stay long enough

Hold time is how long the data must stay steady after the clock edge. It is broken by a path that is too fast, and the clock period cannot help you.

This is the check that surprises people, so go slowly.

At the capture edge, FF2 is reading the value that has been sitting at its input all cycle. At the same edge, FF1 launches a new value. If that new value races through the logic and reaches FF2 almost immediately, it can overwrite the old one while FF2 is still reading it.

Think of it like this

Someone is copying a number off a whiteboard. You may rub it out and write the next one, but not while they are still reading. Hold time is how long you must wait before you are allowed to wipe.

The short path

Take a path with just one gate in it, and use the fastest numbers, because this check is about data being too early.

A short path with only one gate, where the new data arrives almost immediately FF1 D Q t_cq 0.12 one NAND2 0.05 ns FF2 D Q clk period 10.00 ns new data arrives 0.17 ns after the edge
Figure 1.5 - The same clock edge launches at FF1 and captures at FF2. Data leaves FF1 after 0.12 ns at its fastest, and one NAND2 adds 0.05 ns. So the new value is at FF2 only 0.17 ns after the edge, and FF2 needs its old value undisturbed for 0.12 ns.
  1. The new data arrives 0.17 ns after the capture edge.
  2. FF2 needs the old value held for 0.12 ns after that edge.
  3. Hold slack = 0.17 - 0.12 = +0.05 ns. It passes, but only just.
Hold slack, in words hold slack = (how late the new data arrives) - (how long the old value must be held)

Notice what is missing from that sum: the clock period. Both ends of a hold check are the same edge, so the period cancels out.

Break it with a little clock skew

Now suppose the clock reaches FF2 0.15 ns later than it reaches FF1 - one extra buffer in the clock tree. That is clock skew, and Volume 06 is devoted to it.

FF2's edge now happens 0.15 ns later, so the new data is no longer 0.17 ns late relative to that edge. It is 0.02 ns early:

The check Value
New data arrives after FF1's edge 0.17 ns
FF2's edge is later by 0.15 ns
Hold requirement at FF2 0.12 ns
Hold slack -0.10 ns

The path now fails. And here is the point of the whole sub-module:

Clock period Frequency Hold slack
10 ns 100.0 MHz -0.10 ns
100 ns 10.0 MHz -0.10 ns
1000 ns 1.0 MHz -0.10 ns

A thousand times slower, and the violation has not moved a picosecond.

How it is fixed

By making the data slower on purpose. Add two small buffers to the data path, worth 0.25 ns even at their fastest, and the arrival moves from 0.17 ns to 0.42 ns:

hold slack = 0.42 - 0.15 - 0.12 = +0.15 ns, and the path passes.

Buffers that do nothing but waste time

Place-and-route tools insert these automatically, in their thousands. If you ever open a netlist and find a chain of buffers driving nothing but the next buffer, you are looking at a hold fix. Volume 14 does this properly.

Common mistake

Trying to fix a hold violation by lowering the clock frequency. It is the one failure that a slower clock cannot touch. Hold is fixed in the data path or in the clock tree, never in the period.

Quick check

A hold violation of -0.08 ns is found on a design running at 500 MHz. The team drops the clock to 50 MHz. What happens to the violation?

Show the answer

Answer: A. The hold check compares the data with the edge that launched it, so the clock period never appears in the sum. Changing the frequency changes setup slack only. The hold violation is unchanged at -0.08 ns.

1.4 What happens when timing fails

A setup failure gives you a chip that works when it is clocked slower. A hold failure gives you a chip that never works at all. That difference decides how seriously each one is taken.

What the flip-flop actually does

When data changes too close to the clock edge, a flip-flop has three ways to disappoint you.

  1. It captures the old value. The most common outcome, and the quietest. The design carries on with a stale number.
  2. It captures the new value. Sometimes that is even correct, which is worse: the bug appears only on some chips.
  3. It goes metastable. The output sits between 0 and 1 for a while before settling at random. Anything reading it may disagree about what it saw.

None of these prints a message. They produce wrong answers at full speed.

The cost of each failure

Setup failure Hold failure
Cause The data was too slow The data was too fast
Depends on the clock period? Yes No
Fix in the design Shorter logic, more pipelining Add delay, or mend the clock tree
Fix without changing the design Run the chip slower There isn't one
What ships A part sold at a lower speed grade Scrap

That last row is the whole reason hold checks are treated as sacred. A chip with a setup problem is sold as a slower part, and the customer may never know. A chip with a hold problem is a paperweight in every temperature, at every voltage, at every speed.

Binning, in one example

Our good path stops at 212.8 MHz. The version with the 32-bit adder stops at 93.6 MHz. If that one path is in a chip meant for 200 MHz, the chip is not a 200 MHz part. It is a 93.6 MHz part, because the slowest path decides.

Remember

One path sets the speed of the entire chip. Not the average path, not most of them - the worst one. That is why timing reports are always sorted by slack, worst first.

Quick check

Which failure can be made to go away by selling the chip at a lower clock frequency?

Show the answer

Answer: D. A setup failure is about the data being too slow for the period, so a longer period fixes it - the chip is sold at a lower speed grade. A hold failure does not involve the period at all, so a slower clock changes nothing.

1.5 STA versus simulation

Simulation checks the patterns you thought of. Static timing analysis checks every path, whether you thought of it or not.

Both are needed, and they answer different questions. Simulation asks whether the logic is right. Timing analysis asks whether the logic is fast enough.

Why you cannot simulate your way to timing closure

Imagine proving that a 32-bit adder is fast enough by trying inputs. It has 64 input bits, so there are 264 patterns:

The size of the job 264 = 18,446,744,073,709,551,616 patterns

At a billion patterns a second, that is about 585 years for one adder. And the worst case might need two patterns in a row, not one.

Static timing analysis looks at the same adder once. It never asks what the inputs are: it takes the longest delay through the logic and the shortest, and checks both.

What each one is good at

Static timing analysis Gate-level simulation
What it checks Every path in the design The paths your test happens to exercise
Needs test patterns? No Yes, and their quality decides everything
Time for a whole chip Minutes to hours Days, and still not complete
Finds a slow path nobody tested Yes Only by luck
Says whether the logic is correct No Yes
Handles asynchronous inputs well No - they need their own checks Partly
In plain words

The short version: timing analysis is a proof about delay, and simulation is a test of behaviour. Neither replaces the other, and a chip that skips either one does not tape out.

What timing analysis needs from you

It is only as good as what it is told. If you forget to tell the tool about a clock, the paths on that clock are simply not checked, and the report still says zero violations.

That is why Volume 13 spends a whole volume on constraints, and why a real flow ends with a command that reports what was not checked.

Common mistake

Reading "no violations" as "the timing is fine". It means "nothing I was asked to check failed". Always look at how many paths were checked, and at what the tool says it could not analyse.

Quick check

A design passes gate-level simulation with a large set of tests. What does that tell you about its timing?

Show the answer

Answer: B. Simulation can only exercise the paths its patterns reach. A slow path nobody tested stays hidden, which is exactly why static timing analysis - which checks all of them, with no patterns - is a separate, compulsory step.

What you learned

Key words from this volume

Every word below has a plain-English entry in the glossary.

Practice

Practice 1

Slack, both ways

A path's data is ready 7.40 ns after the launch edge. The clock period is 8 ns and the capture flip-flop's setup time is 0.20 ns. What is the setup slack? What is the fastest clock this path could take?

Show the solution

The deadline is 8.00 - 0.20 = 7.80 ns. The data arrives at 7.40 ns, so the slack is 7.80 - 7.40 = +0.40 ns.

For the fastest clock, set the slack to zero. The period must be at least 7.40 + 0.20 = 7.60 ns, which is 1000 / 7.60 = 131.6 MHz.

Practice 2

Which check broke?

A board runs correctly at 50 MHz and at 100 MHz, but produces wrong results at 150 MHz. Which check is failing, and what would you look at first?

Show the solution

It is a setup failure. Only setup depends on the clock period, so a fault that appears when the clock speeds up must be data arriving too late.

The first thing to look at is the worst-slack path in the timing report at 150 MHz. The slack will be negative by roughly the amount the period shrank past the path's limit. A hold problem would have broken the board at 50 MHz as well.

Practice 3

The buffer that fixes nothing

An engineer has a hold violation of -0.10 ns. They add a buffer of 0.25 ns to the clock path going to the capture flip-flop, rather than to the data path. What happens?

Show the solution

It gets worse. Delaying the capture clock is exactly what caused the violation in this volume: the capture edge moves later, so the new data looks even earlier relative to it. The slack would go from -0.10 ns to about -0.35 ns.

The two fixes that work are adding delay to the data path, which is what tools do automatically, or making the capture clock arrive earlier by rebalancing the clock tree.

Practice 4

Reading a speed grade

A chip has three paths, limited to 212.8 MHz, 180 MHz and 93.6 MHz. What is the fastest clock the chip can be sold at, and what happens if you fix only the 93.6 MHz path?

Show the solution

The chip runs at 93.6 MHz, because the slowest path decides. Nothing else matters until that path is fixed.

Fix it, and the limit becomes 180 MHz - the next-worst path takes over. This is why timing closure feels like whack-a-mole: each fix promotes a new critical path, and the reports are sorted worst-first so you always see the one that matters.

Interview corner

Interview question 1

Why does hold not depend on frequency?

"Explain why a hold violation cannot be fixed by lowering the clock frequency."

Show the solution

"Because both sides of the hold check reference the same clock edge. The data is launched by an edge, and the check asks whether it reaches the capture flop before that same edge is finished with. It is one clock-to-Q plus the logic delay, against the hold requirement. The period is not in the equation, so it cancels whatever frequency you pick.

Setup is the opposite: it compares the launch edge with the next one, so the period is right there in the sum. That is why a setup violation can be sold as a slower part, and a hold violation cannot be sold at all."

Interview question 2

Setup failure or hold failure?

"A prototype works at room temperature but fails when the lab heats up. Which check would you suspect, and why?"

Show the solution

"Setup, most likely. Heat makes transistors slower, so the data path gets slower while the period stays the same, and a path with little slack tips over. A hold problem would be the other way round: it shows up when the chip is cold and fast, and it would not care about the clock speed.

The first thing I would do is look at the worst setup slack at the slow corner - hot, low voltage - and see how close it was. If it was a handful of picoseconds, the corner was simply too tight."

Volume 02 goes back a step and asks where all these delay numbers come from. What makes a gate slow, what does a wire cost, and how does a library store it all?