Volume 06 Intermediate 5 sub-modules ~20 min read

The Clock in Detail

Every check so far has leaned on the clock, and treated it kindly. A real clock has a duty cycle that may not be 50%, takes time to leave its source and more to cross the chip, reaches no two flip-flops at quite the same moment, and wobbles from cycle to cycle. This volume takes each of those in turn, and puts a number on what it does to setup and hold.

You will learn
  • How a clock waveform is described, and why the duty cycle matters to half-cycle paths
  • The difference between source latency and network latency, and when each one counts
  • Local, global and useful skew, and how skew can lend time between paths
  • What jitter is, and how clock uncertainty is built before and after the clock tree
  • Why ideal and propagated clocks give different slacks, and when to trust each
You need

6.1 Period, duty cycle and waveform

A clock is described by three numbers: its period, when it rises, and when it falls. The last two set the duty cycle, and any path that uses both edges of the clock lives or dies by it.

A timing tool knows a clock only by what it is told. In the constraint language the course meets in Volume 13, one line says it all:


create_clock -name clk -period 10 -waveform {0 5} [get_ports clk]

This says: a clock called clk, 10 ns long, rising at 0 and falling at 5, arriving on the port clk. The waveform gives the edge times within one period, and every later edge is a whole period on.

Clock Rises at Falls at
10 ns, 50% duty, waveform {0 5} 0, 10, 20 5, 15
10 ns, 40% duty, waveform {0 4} 0, 10, 20 4, 14
10 ns, 50% duty, waveform {2 7} 2, 12 7, 17
A 50% clock, a 40% clock, a 60% clock and a clock that rises 2 ns late 0 1 2 3 clk clk_40 clk_60 clk_late
Figure 6.1 - Four clocks with the same 10 ns period. The rising edges of the first three line up; only their falling edges move. The last one is the 50% clock shifted by 2 ns - the same shape, a different waveform.

Why the duty cycle matters

A path from a rising-edge flip-flop to a falling-edge flip-flop gets only the time from the rise to the fall. That is the high part of the clock:

Duty cycle Rise-to-fall path gets Fall-to-rise path gets
50% 5.00 ns 5.00 ns
40% 4.00 ns 6.00 ns
60% 6.00 ns 4.00 ns

Take a rise-to-fall path needing 3.70 ns for clock-to-Q and logic, with a 0.12 ns setup time. At 50% duty its slack is 1.18 ns. Let the duty cycle slip to 40% and the slack falls to 0.18 ns. The period never changed.

Common mistake

Assuming a clock is exactly 50% because the constraint says so. A real clock's duty cycle drifts with the PLL, the buffers and the chip's temperature. Any design that uses both edges must be checked at the worst duty cycle it can see, not at the ideal one.

Quick check

An 8 ns clock has a 25% duty cycle: high for a quarter of the period. How much time does a rise-to-fall path get?

Show the answer

Answer: B. The clock rises at 0 and falls a quarter of the period later, at 2 ns. A rise-to-fall path gets exactly that high time: 2 ns. A fall-to-rise path would get the other 6 ns.

6.2 Clock latency: source and network

Clock latency comes in two parts. Source latency is the time from where the clock is made to where the design's clock begins. Network latency is the time across the design's own clock tree.

The clock's journey: from a PLL, through the clock port, across the clock tree to two flip-flops PLL clock port FF1 FF2 edge arrives at 0.85 ns edge arrives at 0.88 ns source latency 0.50 network latency 0.35 or 0.38
Figure 6.2 - Source latency is spent before the clock reaches the design. Network latency is spent inside it, in the clock tree. The times at the right are when the edge reaches each flip-flop. Each is 0.50 ns of source latency plus 0.35 or 0.38 ns of tree, as in Volume 04.

For paths inside the design, the source part cancels

Both flip-flops of a register-to-register path get their clock from the same source. So the source latency appears on the arrival side and the required side alike, and drops out. Only the network latency to each flip-flop - and so the skew - is left.

For paths to and from the outside world, it is different

An input path starts at a flip-flop in another chip. That flip-flop shares the clock source, so the source latency still cancels. But it does not share our clock tree. So the network latency to our capture flip-flop is left over, and it helps: the capture edge arrives later.

An output path is the mirror image: our launch flip-flop's network latency makes the data leave later, and it hurts.

Clock latency Register to register Input path Output path
None 1.81 ns 0.71 ns 1.79 ns
Source 1.00, network 0.60 to every flip-flop 1.81 ns 1.31 ns 1.19 ns
Source 1.00, network 0.60 launch, 0.70 capture 1.91 ns 1.41 ns 1.19 ns

All at 5 ns. The input path has a 3.00 ns input delay and 1.20 ns of logic. The output path has 0.21 + 1.00 ns, and a 2.00 ns output delay. The source latency changed nothing anywhere. The network latency moved the input and output paths by exactly its own size.

Remember

Source latency is shared by everyone on the clock, so it cancels. Network latency belongs to the flip-flops inside the design, so it matters wherever only one end of a path is inside.

Common mistake

Forgetting that a long clock tree eats into the output paths. A chip whose clock tree is 1.5 ns deep sends its outputs 1.5 ns late, and the board designer's timing budget did not know that. This is why fast interfaces use special clocking - Volume 10 comes back to it.

Quick check

A register-to-register path has 0.40 ns of slack. The clock's source latency is then increased by 1 ns. What happens to the slack?

Show the answer

Answer: D. Both flip-flops receive the clock from the same source, so the extra 1 ns delays the launch and the capture equally. It appears on both sides of the sum and cancels.

6.3 Skew: positive, negative and useful

Skew is the difference in clock arrival between two flip-flops. It helps setup when the capture clock is later, and hurts hold by the same amount. Designers can even add it on purpose.

Local skew and global skew

Take four flip-flops, with clock latencies of 0.82, 0.85, 0.91 and 0.88 ns. Two numbers describe their skew:

Measure What it compares Value here
Global skew The latest and earliest arrival anywhere 0.91 - 0.82 = 0.09 ns
Local skew, FF1 to FF2 Capture minus launch, on one path +0.03 ns
Local skew, FF2 to FF3 +0.06 ns
Local skew, FF3 to FF4 -0.03 ns
Local skew, FF4 to FF1 -0.06 ns

Global skew describes the clock tree. Local skew is what a path actually feels. Only local skew enters a timing check, and it has a sign.

Positive and negative skew

  1. Positive skew: the capture flip-flop gets the clock later. Setup gains; hold loses.
  2. Negative skew: the capture flip-flop gets it earlier. Setup loses; hold gains.

Volume 05 drew the window of skew where both checks pass. Every pair of flip-flops has one.

Useful skew: lending time

Two stages in a row, on a 5 ns clock. Stage A, from FF1 to FF2, fails setup by 0.30 ns. Stage B, from FF2 to FF3, has 0.80 ns to spare. FF2 sits in the middle of both.

Delay FF2's clock by 0.40 ns. Stage A gets 0.40 ns more to reach FF2. Stage B loses 0.40 ns, because FF2 now launches later.

Clock at FF2 A: setup A: hold B: setup B: hold
Balanced -0.30 ns 0.68 ns 0.80 ns 0.53 ns
0.40 ns late +0.10 ns 0.28 ns +0.40 ns 0.93 ns

Both stages now pass setup. Stage A's hold slack dropped by 0.40 ns but is still comfortable; stage B's hold improved. This is useful skew, and clock tree tools use it on purpose.

Common mistake

Borrowing through useful skew without checking hold. The borrowed time comes out of the capturing flip-flop's hold margin. With a short path into FF2, the same 0.40 ns would have broken hold.

Quick check

A 4 ns path fails setup by 0.20 ns. Its capture clock is delayed by 0.30 ns. What happens?

Show the answer

Answer: C. A later capture clock adds its delay to the setup slack: -0.20 + 0.30 = +0.10 ns. The hold check loses the same 0.30 ns. In this example hold still passes, with 0.17 ns left.

6.4 Jitter and uncertainty

Jitter is the clock edge wandering from cycle to cycle. Nobody can remove it in the design, so a margin is taken for it - and for anything else not yet known - as clock uncertainty.

A clock comes from a PLL or an oscillator, and neither is perfect. Each edge lands a few picoseconds early or late compared with a perfect clock. From one cycle to the next the period is never quite the same.

A clock whose rising edges wander a little from where a perfect clock would put them jitter one period, on average
Figure 6.3 - Jitter. The square wave shows where a perfect clock would put each edge. The thin gold lines show where the real edge actually lands on different cycles. Setup has to survive the worst case: one edge late and the next one early.

Building the uncertainty

Clock uncertainty is a single number, but it is built from parts. Before the clock tree exists, it also has to stand in for skew nobody knows yet.

Before the clock tree After the clock tree
Setup uncertainty jitter 0.05 + skew estimate 0.10 + margin 0.03 = 0.18 ns jitter 0.05 + margin 0.03 = 0.08 ns
Hold uncertainty skew estimate 0.10 + margin 0.03 = 0.13 ns margin 0.03 ns

Two things to notice. After the tree is built, the skew estimate goes, because the real skew is now in the clock latencies. And hold carries no jitter at all: its launch and capture are the same edge, so that edge cannot differ from itself.

What it costs

At 5 ns, the Volume 04 logic has 1.63 ns of setup slack with the early uncertainty of 0.18 ns. With the later 0.08 ns it has 1.73 ns. Jitter matters more as clocks get faster: 50 ps of jitter is 5% of a 1 GHz cycle, but only 0.5% of a 100 MHz one.

Common mistake

Leaving the pre-clock-tree uncertainty in place after the tree is built. The skew is then counted twice - once in the real latencies and once in the margin - and the design looks worse than it is.

Quick check

A team uses 40 ps of jitter, a 120 ps skew estimate and a 20 ps margin before the clock tree is built. What setup uncertainty should they set?

Show the answer

Answer: A. Before the clock tree, setup uncertainty is jitter plus the skew estimate plus the margin: 40 + 120 + 20 = 180 ps, or 0.18 ns. After the tree is built, it would drop to 40 + 20 = 60 ps.

6.5 Ideal versus propagated clocks

An ideal clock reaches every flip-flop at once; a propagated clock uses the real delays through the built tree. Early in a project you only have the first, and it is a guess. The second is the truth.

The clock tree is built by clock tree synthesis, after the cells have been placed. Until then there is no tree to measure. So the tool treats the clock as ideal, and leans on the uncertainty to cover the skew it cannot see.


# before clock tree synthesis: ideal clock, generous margin
set_clock_uncertainty -setup 0.18 [get_clocks clk]
set_clock_uncertainty -hold  0.13 [get_clocks clk]

# after clock tree synthesis: measure the real tree
set_propagated_clock [all_clocks]
set_clock_uncertainty -setup 0.08 [get_clocks clk]
set_clock_uncertainty -hold  0.03 [get_clocks clk]

The same two paths, both ways

Path Check Ideal clock Propagated clock
The adder path (Volume 04) Setup 1.63 ns 1.76 ns
The short path (Volume 05) Hold +0.03 ns -0.01 ns

The adder path looked worse with the ideal clock, because the generous margin was pessimistic for it. The short path looked better, because the guessed 0.10 ns of skew was smaller than the real 0.14 ns. Its hold violation only appeared once the real tree was measured.

In plain words

An ideal clock gives you a forecast. A propagated clock gives you the weather. Plan with the first, but sign off only with the second.

Remember

The first timing run after clock tree synthesis is the one that shows the real hold picture. It is normal for it to reveal new violations; the tools fix most of them by inserting delay cells.

Quick check

Why does hold uncertainty drop from 0.13 ns to 0.03 ns after clock tree synthesis?

Show the answer

Answer: B. Before the tree, 0.10 ns of the uncertainty stood in for skew nobody knew. Once the tree exists, the real latency to every flip-flop is measured, so that estimate would count the skew twice. Only the small margin remains.

What you learned

Key words from this volume

Every word below has a plain-English entry in the glossary.

Practice

Practice 1

Global skew

Four flip-flops receive the clock at 1.10, 1.25, 1.18 and 1.31 ns. What is the global skew? And what is the local skew on a path from the 1.31 ns flip-flop to the 1.10 ns one?

Show the solution

Global skew is the latest minus the earliest: 1.31 - 1.10 = 0.21 ns.

Local skew is capture minus launch: 1.10 - 1.31 = -0.21 ns. It is negative, so that path loses 0.21 ns of setup and gains 0.21 ns of hold.

Practice 2

A duty-cycle trap

A design launches data on the rising edge and captures it on the falling edge of a 10 ns clock. It was signed off at 50% duty with 1.18 ns of slack. The clock generator is replaced by one with a 40% duty cycle. Does the path still pass?

Show the solution

The path now gets 4.00 ns instead of 5.00 ns, so its slack drops by 1.00 ns: 1.18 - 1.00 = 0.18 ns. It still passes, but with little left.

A 35% duty cycle would have broken it. Designs that use both clock edges state the duty cycle they need, and the clock source must guarantee it.

Practice 3

Useful skew, the other way

In the useful-skew example, what would happen if FF2's clock were delayed by 0.90 ns instead of 0.40 ns?

Show the solution

Stage A would gain 0.90 ns of setup: -0.30 + 0.90 = +0.60 ns. But stage B would lose 0.90 ns: 0.80 - 0.90 = -0.10 ns, a new violation. And stage A's hold slack would fall from 0.68 to -0.22 ns.

Useful skew moves slack between neighbours; it does not create it. Borrow too much and the lender fails instead.

Interview corner

Interview question 1

Skew and jitter

"What is the difference between clock skew and clock jitter, and how does each enter STA?"

Show the solution

"Skew is spatial: the same edge reaches different flip-flops at different times. It is fixed for a given clock tree, so after clock tree synthesis it is measured directly through propagated clock latencies. Jitter is temporal: the edge at one place moves from cycle to cycle. It cannot be calculated from the netlist, so it is added as clock uncertainty on setup checks.

Hold checks normally carry no jitter, because the launch and capture are the same edge. Before clock tree synthesis, the uncertainty also includes an estimate of skew; that estimate is removed once the clock is propagated."

Interview question 2

Why does source latency not matter?

"A clock has 2 ns of source latency from an off-chip PLL. Does it affect your internal paths?"

Show the solution

"No. Every flip-flop on the clock sees the same source latency, so it shifts the launch and the capture edges equally and cancels in every register-to-register check. It would only matter against something that does not share that source - for example an interface timed from a different clock. The network latency inside the chip is what creates skew, and that is what I would look at."

Volume 07 takes every clocking case in turn. It covers paths between rising and falling edges, divided and generated clocks, and clocks with awkward ratios. Then come clocks that have nothing to do with each other, and clocks that pass through multiplexers and gates.