The Clock in Detail
Every check so far has leaned on the clock, and treated it kindly. A real clock has a duty cycle that may not be 50%, takes time to leave its source and more to cross the chip, reaches no two flip-flops at quite the same moment, and wobbles from cycle to cycle. This volume takes each of those in turn, and puts a number on what it does to setup and hold.
- How a clock waveform is described, and why the duty cycle matters to half-cycle paths
- The difference between source latency and network latency, and when each one counts
- Local, global and useful skew, and how skew can lend time between paths
- What jitter is, and how clock uncertainty is built before and after the clock tree
- Why ideal and propagated clocks give different slacks, and when to trust each
6.1 Period, duty cycle and waveform
A clock is described by three numbers: its period, when it rises, and when it falls. The last two set the duty cycle, and any path that uses both edges of the clock lives or dies by it.
A timing tool knows a clock only by what it is told. In the constraint language the course meets in Volume 13, one line says it all:
create_clock -name clk -period 10 -waveform {0 5} [get_ports clk]
This says: a clock called clk, 10 ns long, rising at 0 and falling at 5, arriving on the port clk. The waveform gives the edge times within one period, and every later edge is a whole period on.
| Clock | Rises at | Falls at |
|---|---|---|
| 10 ns, 50% duty, waveform {0 5} | 0, 10, 20 | 5, 15 |
| 10 ns, 40% duty, waveform {0 4} | 0, 10, 20 | 4, 14 |
| 10 ns, 50% duty, waveform {2 7} | 2, 12 | 7, 17 |
Why the duty cycle matters
A path from a rising-edge flip-flop to a falling-edge flip-flop gets only the time from the rise to the fall. That is the high part of the clock:
| Duty cycle | Rise-to-fall path gets | Fall-to-rise path gets |
|---|---|---|
| 50% | 5.00 ns | 5.00 ns |
| 40% | 4.00 ns | 6.00 ns |
| 60% | 6.00 ns | 4.00 ns |
Take a rise-to-fall path needing 3.70 ns for clock-to-Q and logic, with a 0.12 ns setup time. At 50% duty its slack is 1.18 ns. Let the duty cycle slip to 40% and the slack falls to 0.18 ns. The period never changed.
Assuming a clock is exactly 50% because the constraint says so. A real clock's duty cycle drifts with the PLL, the buffers and the chip's temperature. Any design that uses both edges must be checked at the worst duty cycle it can see, not at the ideal one.
An 8 ns clock has a 25% duty cycle: high for a quarter of the period. How much time does a rise-to-fall path get?
Show the answer
Answer: B. The clock rises at 0 and falls a quarter of the period later, at 2 ns. A rise-to-fall path gets exactly that high time: 2 ns. A fall-to-rise path would get the other 6 ns.
6.2 Clock latency: source and network
Clock latency comes in two parts. Source latency is the time from where the clock is made to where the design's clock begins. Network latency is the time across the design's own clock tree.
For paths inside the design, the source part cancels
Both flip-flops of a register-to-register path get their clock from the same source. So the source latency appears on the arrival side and the required side alike, and drops out. Only the network latency to each flip-flop - and so the skew - is left.
For paths to and from the outside world, it is different
An input path starts at a flip-flop in another chip. That flip-flop shares the clock source, so the source latency still cancels. But it does not share our clock tree. So the network latency to our capture flip-flop is left over, and it helps: the capture edge arrives later.
An output path is the mirror image: our launch flip-flop's network latency makes the data leave later, and it hurts.
| Clock latency | Register to register | Input path | Output path |
|---|---|---|---|
| None | 1.81 ns | 0.71 ns | 1.79 ns |
| Source 1.00, network 0.60 to every flip-flop | 1.81 ns | 1.31 ns | 1.19 ns |
| Source 1.00, network 0.60 launch, 0.70 capture | 1.91 ns | 1.41 ns | 1.19 ns |
All at 5 ns. The input path has a 3.00 ns input delay and 1.20 ns of logic. The output path has 0.21 + 1.00 ns, and a 2.00 ns output delay. The source latency changed nothing anywhere. The network latency moved the input and output paths by exactly its own size.
Source latency is shared by everyone on the clock, so it cancels. Network latency belongs to the flip-flops inside the design, so it matters wherever only one end of a path is inside.
Forgetting that a long clock tree eats into the output paths. A chip whose clock tree is 1.5 ns deep sends its outputs 1.5 ns late, and the board designer's timing budget did not know that. This is why fast interfaces use special clocking - Volume 10 comes back to it.
A register-to-register path has 0.40 ns of slack. The clock's source latency is then increased by 1 ns. What happens to the slack?
Show the answer
Answer: D. Both flip-flops receive the clock from the same source, so the extra 1 ns delays the launch and the capture equally. It appears on both sides of the sum and cancels.
6.3 Skew: positive, negative and useful
Skew is the difference in clock arrival between two flip-flops. It helps setup when the capture clock is later, and hurts hold by the same amount. Designers can even add it on purpose.
Local skew and global skew
Take four flip-flops, with clock latencies of 0.82, 0.85, 0.91 and 0.88 ns. Two numbers describe their skew:
| Measure | What it compares | Value here |
|---|---|---|
| Global skew | The latest and earliest arrival anywhere | 0.91 - 0.82 = 0.09 ns |
| Local skew, FF1 to FF2 | Capture minus launch, on one path | +0.03 ns |
| Local skew, FF2 to FF3 | +0.06 ns | |
| Local skew, FF3 to FF4 | -0.03 ns | |
| Local skew, FF4 to FF1 | -0.06 ns |
Global skew describes the clock tree. Local skew is what a path actually feels. Only local skew enters a timing check, and it has a sign.
Positive and negative skew
- Positive skew: the capture flip-flop gets the clock later. Setup gains; hold loses.
- Negative skew: the capture flip-flop gets it earlier. Setup loses; hold gains.
Volume 05 drew the window of skew where both checks pass. Every pair of flip-flops has one.
Useful skew: lending time
Two stages in a row, on a 5 ns clock. Stage A, from FF1 to FF2, fails setup by 0.30 ns. Stage B, from FF2 to FF3, has 0.80 ns to spare. FF2 sits in the middle of both.
Delay FF2's clock by 0.40 ns. Stage A gets 0.40 ns more to reach FF2. Stage B loses 0.40 ns, because FF2 now launches later.
| Clock at FF2 | A: setup | A: hold | B: setup | B: hold |
|---|---|---|---|---|
| Balanced | -0.30 ns | 0.68 ns | 0.80 ns | 0.53 ns |
| 0.40 ns late | +0.10 ns | 0.28 ns | +0.40 ns | 0.93 ns |
Both stages now pass setup. Stage A's hold slack dropped by 0.40 ns but is still comfortable; stage B's hold improved. This is useful skew, and clock tree tools use it on purpose.
Borrowing through useful skew without checking hold. The borrowed time comes out of the capturing flip-flop's hold margin. With a short path into FF2, the same 0.40 ns would have broken hold.
A 4 ns path fails setup by 0.20 ns. Its capture clock is delayed by 0.30 ns. What happens?
Show the answer
Answer: C. A later capture clock adds its delay to the setup slack: -0.20 + 0.30 = +0.10 ns. The hold check loses the same 0.30 ns. In this example hold still passes, with 0.17 ns left.
6.4 Jitter and uncertainty
Jitter is the clock edge wandering from cycle to cycle. Nobody can remove it in the design, so a margin is taken for it - and for anything else not yet known - as clock uncertainty.
A clock comes from a PLL or an oscillator, and neither is perfect. Each edge lands a few picoseconds early or late compared with a perfect clock. From one cycle to the next the period is never quite the same.
Building the uncertainty
Clock uncertainty is a single number, but it is built from parts. Before the clock tree exists, it also has to stand in for skew nobody knows yet.
| Before the clock tree | After the clock tree | |
|---|---|---|
| Setup uncertainty | jitter 0.05 + skew estimate 0.10 + margin 0.03 = 0.18 ns | jitter 0.05 + margin 0.03 = 0.08 ns |
| Hold uncertainty | skew estimate 0.10 + margin 0.03 = 0.13 ns | margin 0.03 ns |
Two things to notice. After the tree is built, the skew estimate goes, because the real skew is now in the clock latencies. And hold carries no jitter at all: its launch and capture are the same edge, so that edge cannot differ from itself.
What it costs
At 5 ns, the Volume 04 logic has 1.63 ns of setup slack with the early uncertainty of 0.18 ns. With the later 0.08 ns it has 1.73 ns. Jitter matters more as clocks get faster: 50 ps of jitter is 5% of a 1 GHz cycle, but only 0.5% of a 100 MHz one.
Leaving the pre-clock-tree uncertainty in place after the tree is built. The skew is then counted twice - once in the real latencies and once in the margin - and the design looks worse than it is.
A team uses 40 ps of jitter, a 120 ps skew estimate and a 20 ps margin before the clock tree is built. What setup uncertainty should they set?
Show the answer
Answer: A. Before the clock tree, setup uncertainty is jitter plus the skew estimate plus the margin: 40 + 120 + 20 = 180 ps, or 0.18 ns. After the tree is built, it would drop to 40 + 20 = 60 ps.
6.5 Ideal versus propagated clocks
An ideal clock reaches every flip-flop at once; a propagated clock uses the real delays through the built tree. Early in a project you only have the first, and it is a guess. The second is the truth.
The clock tree is built by clock tree synthesis, after the cells have been placed. Until then there is no tree to measure. So the tool treats the clock as ideal, and leans on the uncertainty to cover the skew it cannot see.
# before clock tree synthesis: ideal clock, generous margin
set_clock_uncertainty -setup 0.18 [get_clocks clk]
set_clock_uncertainty -hold 0.13 [get_clocks clk]
# after clock tree synthesis: measure the real tree
set_propagated_clock [all_clocks]
set_clock_uncertainty -setup 0.08 [get_clocks clk]
set_clock_uncertainty -hold 0.03 [get_clocks clk]
The same two paths, both ways
| Path | Check | Ideal clock | Propagated clock |
|---|---|---|---|
| The adder path (Volume 04) | Setup | 1.63 ns | 1.76 ns |
| The short path (Volume 05) | Hold | +0.03 ns | -0.01 ns |
The adder path looked worse with the ideal clock, because the generous margin was pessimistic for it. The short path looked better, because the guessed 0.10 ns of skew was smaller than the real 0.14 ns. Its hold violation only appeared once the real tree was measured.
An ideal clock gives you a forecast. A propagated clock gives you the weather. Plan with the first, but sign off only with the second.
The first timing run after clock tree synthesis is the one that shows the real hold picture. It is normal for it to reveal new violations; the tools fix most of them by inserting delay cells.
Why does hold uncertainty drop from 0.13 ns to 0.03 ns after clock tree synthesis?
Show the answer
Answer: B. Before the tree, 0.10 ns of the uncertainty stood in for skew nobody knew. Once the tree exists, the real latency to every flip-flop is measured, so that estimate would count the skew twice. Only the small margin remains.
What you learned
- A clock is described by its period and its waveform; the duty cycle sets the time for half-cycle paths.
- Source latency is shared, so it cancels; network latency matters wherever a path leaves the design.
- Global skew describes a clock tree; local skew, with its sign, is what a path feels.
- Useful skew lends setup time from one stage to the next, at a cost in hold.
- Jitter is cycle-to-cycle wander, covered by clock uncertainty, and it does not apply to hold.
- Before the clock tree, uncertainty also covers unknown skew; afterwards, the real latencies do.
- Ideal clocks are for planning; propagated clocks are for sign-off.
Key words from this volume
Every word below has a plain-English entry in the glossary.
- Duty cycle
- Source latency
- Network latency
- Useful skew
- Jitter
- PLL (phase-locked loop)
- Clock uncertainty
- Ideal clock
- Propagated clock
- Clock tree synthesis (CTS)
Practice
Global skew
Four flip-flops receive the clock at 1.10, 1.25, 1.18 and 1.31 ns. What is the global skew? And what is the local skew on a path from the 1.31 ns flip-flop to the 1.10 ns one?
Show the solution
Global skew is the latest minus the earliest: 1.31 - 1.10 = 0.21 ns.
Local skew is capture minus launch: 1.10 - 1.31 = -0.21 ns. It is negative, so that path loses 0.21 ns of setup and gains 0.21 ns of hold.
A duty-cycle trap
A design launches data on the rising edge and captures it on the falling edge of a 10 ns clock. It was signed off at 50% duty with 1.18 ns of slack. The clock generator is replaced by one with a 40% duty cycle. Does the path still pass?
Show the solution
The path now gets 4.00 ns instead of 5.00 ns, so its slack drops by 1.00 ns: 1.18 - 1.00 = 0.18 ns. It still passes, but with little left.
A 35% duty cycle would have broken it. Designs that use both clock edges state the duty cycle they need, and the clock source must guarantee it.
Useful skew, the other way
In the useful-skew example, what would happen if FF2's clock were delayed by 0.90 ns instead of 0.40 ns?
Show the solution
Stage A would gain 0.90 ns of setup: -0.30 + 0.90 = +0.60 ns. But stage B would lose 0.90 ns: 0.80 - 0.90 = -0.10 ns, a new violation. And stage A's hold slack would fall from 0.68 to -0.22 ns.
Useful skew moves slack between neighbours; it does not create it. Borrow too much and the lender fails instead.
Interview corner
Skew and jitter
"What is the difference between clock skew and clock jitter, and how does each enter STA?"
Show the solution
"Skew is spatial: the same edge reaches different flip-flops at different times. It is fixed for a given clock tree, so after clock tree synthesis it is measured directly through propagated clock latencies. Jitter is temporal: the edge at one place moves from cycle to cycle. It cannot be calculated from the netlist, so it is added as clock uncertainty on setup checks.
Hold checks normally carry no jitter, because the launch and capture are the same edge. Before clock tree synthesis, the uncertainty also includes an estimate of skew; that estimate is removed once the clock is propagated."
Why does source latency not matter?
"A clock has 2 ns of source latency from an off-chip PLL. Does it affect your internal paths?"
Show the solution
"No. Every flip-flop on the clock sees the same source latency, so it shifts the launch and the capture edges equally and cancels in every register-to-register check. It would only matter against something that does not share that source - for example an interface timed from a different clock. The network latency inside the chip is what creates skew, and that is what I would look at."
Volume 07 takes every clocking case in turn. It covers paths between rising and falling edges, divided and generated clocks, and clocks with awkward ratios. Then come clocks that have nothing to do with each other, and clocks that pass through multiplexers and gates.