Volume 07 Intermediate 6 sub-modules ~20 min read

Every Clocking Case

Up to now every path ran from a rising edge to the next rising edge of the same clock. Real designs mix falling edges, divided clocks, clocks at awkward ratios, clocks with no relationship at all, and clocks that pass through gates and switches. One rule decides which edges are compared in every one of these cases. This volume applies it to each, with numbers.

You will learn
  • How much time a half-cycle path gets, and why its hold check becomes easy
  • How inverted and generated clocks change the edges a tool compares
  • Why the tightest setup check between two related clocks is the GCD of their periods
  • Why unrelated clocks must be put in separate clock groups, and what that does and does not fix
  • How clock multiplexers and clock gates are checked
You need
  • Volume 06, for clock waveforms, latency and skew

7.1 Half-cycle paths, rise to fall

One rule chooses the edges for every check in this volume. For setup, take each launch edge and find the first capture edge after it; the closest such pair is checked. For hold, check one capture edge earlier than that.

Volume 03 met the rule in its simplest form. Here it meets a half-cycle path: launched on a rising edge and captured on a falling one, or the other way round.

Take one path on a 10 ns, 50% clock. Its clock-to-Q is 0.30 ns (0.20 fastest) and its logic 3.40 ns (0.50 fastest). Setup is 0.12 ns and hold 0.05 ns. Launch and capture it on different edges:

Launch on Capture on Setup edges Setup slack Hold edges Hold slack
Rise Rise 0 then 10 6.18 ns 0 then 0 0.65 ns
Rise Fall 0 then 5 1.18 ns 0 then -5 5.65 ns
Fall Rise 5 then 10 1.18 ns 5 then 0 5.65 ns

The half-cycle paths lose 5 ns of setup, because they get half a period. But look at hold. The capture edge one period before the setup capture comes half a period before the launch. So the data has a whole half-period to spare, and hold slack jumps from 0.65 to 5.65 ns.

In plain words

A half-cycle path trades setup for hold. Half the time to get there, but no danger of arriving too soon. Designs use this deliberately for paths that are short and fast.

Common mistake

Checking only setup on a design with half-cycle paths, because hold looks easy. It is easy for these paths - but the same flip-flops usually have full-cycle paths too, and those still need a hold check.

Quick check

An 8 ns, 50% clock. A rise-to-fall path has 0.20 ns of clock-to-Q, 3.10 ns of logic and a 0.10 ns setup time. What is its setup slack?

Show the answer

Answer: C. A rise-to-fall path gets half the period: 4.00 ns. Slack = 4.00 - 0.20 - 3.10 - 0.10 = 0.60 ns. The 4.60 answer wrongly gives the path the full 8 ns.

7.2 Inverted and negative-edge clocks

A flip-flop that acts on the falling edge, and a flip-flop whose clock passes through an inverter, are the same thing to a timing tool: both capture on the fall. The inverter just adds its delay to that clock's latency.

There are two ways to build a flip-flop that acts on the falling edge. Use a cell made for it, or put an inverter in front of an ordinary rising-edge flip-flop's clock pin. The tool traces the clock through the inverter, sees the edge flip, and times the path exactly like a falling-edge capture.

A clock and the same clock after an inverter 0 1 2 3 clk clk_n
Figure 7.1 - The inverted clock rises when the original falls. A rising-edge flip-flop on the inverted clock therefore captures at the original clock's falling edges - half a period after each rising edge.

The only difference is the inverter's own delay, which arrives on the capture side. For the rise-to-fall path above, an inverter of 0.06 ns in the capture clock gives:

Setup slack Hold slack
Falling-edge flip-flop 1.18 ns 5.65 ns
Rising-edge flip-flop behind a 0.06 ns inverter 1.24 ns 5.59 ns

The later capture clock helps setup and hurts hold by the inverter's 0.06 ns, as any skew would.

Remember

The tool does not care how a falling edge is made. It follows the clock through every buffer and inverter to each flip-flop, and uses the edge that actually arrives there, and when.

Common mistake

Leaving a clock inverter out of the clock tree's balancing. Its delay is real skew between the flops it feeds and every other flop on the clock, and clock tree tools must see it to compensate.

Quick check

A clock passes through an inverter before reaching a rising-edge flip-flop. Which edge of the original clock does that flip-flop capture on?

Show the answer

Answer: A. After the inverter, the clock's rising edges happen when the original clock falls. The flip-flop still acts on its own rising edge, so it captures at the original clock's falling edges.

7.3 Generated and divided clocks

A generated clock is made inside the design from another clock, such as a divide-by-2 from a flip-flop. Declaring it as generated tells the tool how its edges line up with the master, and adds the generating logic's delay to its latency.

A flip-flop whose output feeds back through an inverter toggles on every rising clock edge, making a clock at half the frequency. From a 5 ns master, the result is a 10 ns clock.

A 5 ns master clock and the divide-by-2 clock made from it 0 1 2 3 4 5 clk clk_div2 launch (clk) capture (clk_div2)
Figure 7.2 - The divided clock changes just after each rising edge of the master, because a flip-flop makes it. The tightest setup pair from the master to the divided clock is marked: launched by the master at 5 ns, captured by the divided clock's rise at 10 ns.

create_clock -name clk -period 5 [get_ports clk]
create_generated_clock -name clk_div2 -source [get_ports clk] -divide_by 2 [get_pins u_div/Q]

Which edges get compared

Direction Every setup pair Tightest setup Hold pair
clk to clk_div2 0->10, 5->10, 10->20, 15->20 5 -> 10: 5.00 ns 0 and 0
clk_div2 to clk 0->5, 10->15, 20->25, 30->35 0 -> 5: 5.00 ns 0 and 0

Both directions get one period of the fast clock. From fast to slow, the fast clock's second launch edge is the one closest to the slow capture. From slow to fast, the fast clock captures one of its own periods after the slow launch.

Why the declaration matters

The divider flip-flop takes 0.25 ns to produce each edge of clk_div2. Declared as a generated clock, that 0.25 ns becomes part of clk_div2's latency. Declared as an unrelated clock with its own create_clock, it is lost.

A 4.30 ns path (clock-to-Q 0.20, setup 0.10) Generated clock Unrelated create_clock
clk to clk_div2 0.65 ns 0.40 ns
clk_div2 to clk 0.15 ns 0.40 ns

One direction looks 0.25 ns worse than the truth; the other looks 0.25 ns better. The optimistic one is the dangerous one: a chip can pass timing on paper and fail on the bench.

Common mistake

Using create_clock on the output of a divider. The tool then treats the divided clock as starting fresh at time zero, with no latency from the divider, and often as unrelated to its master. Every divided, multiplied or muxed clock made inside the design should be a generated clock.

Quick check

A path runs from a 5 ns clock to its divide-by-2 clock. How much time does the setup check allow?

Show the answer

Answer: B. The fast clock launches at 0 and at 5 ns; the slow clock captures at 10 ns. The launch at 5 ns is the closer one, so the check allows 10 - 5 = 5 ns: one period of the fast clock.

7.4 Related clocks with odd ratios

When two clocks come from the same source but at an awkward ratio, their edges only line up now and then. The tightest setup check between them is the greatest common divisor of their periods - which can be far shorter than either period.

Take a 125 MHz clock (8 ns) and a 100 MHz clock (10 ns), both from one PLL, with edges lined up at time zero. They line up again only every 40 ns. In between, the launch and capture edges fall in every possible relationship:

Launch (125 MHz) First capture after it (100 MHz) Gap
0 ns 10 ns 10 ns
8 ns 10 ns 2 ns
16 ns 20 ns 4 ns
24 ns 30 ns 6 ns
32 ns 40 ns 8 ns

The tool checks the worst: 2 ns. A path that would be easy on either clock alone has to fit in 2 ns when it crosses between them.

A 125 MHz clock and a 100 MHz clock over 40 ns, with the tightest pair of edges shaded the tightest pair: 8 ns to 10 ns c125 c100
Figure 7.3 - Each block is 2 ns. The two clocks line up at 0 and again at 40 ns. The 125 MHz clock rises at 8 ns and the 100 MHz clock at 10 ns: that 2 ns gap is the setup check every path between them must meet.

The rule

For two clocks whose edges line up at zero, the tightest setup pair is the greatest common divisor (GCD) of their periods. The engine behind this course checks every case in the table against that rule:

Clocks Periods Tightest setup GCD of the periods
100 MHz to 150 MHz 10.00 / 6.67 ns 3.33 ns 3.33 ns
150 MHz to 100 MHz 6.67 / 10.00 ns 3.33 ns 3.33 ns
125 MHz to 100 MHz 8.00 / 10.00 ns 2.00 ns 2.00 ns
100 MHz to 125 MHz 10.00 / 8.00 ns 2.00 ns 2.00 ns
250 MHz to 166.7 MHz 4.00 / 6.00 ns 2.00 ns 2.00 ns
100 MHz to 50 MHz 10.00 / 20.00 ns 10.00 ns 10.00 ns

In every case the hold check lands at 0: the edges that line up at time zero give the usual same-edge hold check.

Take a 2.30 ns path from the 125 MHz clock to the 100 MHz clock, with 0.20 ns clock-to-Q and 0.10 ns setup. Its slack is 2.00 - 0.20 - 2.30 - 0.10 = -0.60 ns. It fails, although either clock alone gives 8 ns or more.

Common mistake

Assuming a crossing between two synchronous clocks gets the shorter of the two periods. It gets their GCD, which can be much smaller. Pick clock frequencies with a simple ratio - 2:1, 3:2 - when paths must cross between them.

Quick check

Two clocks from the same PLL have periods of 6 ns and 9 ns, lined up at zero. What is the tightest setup check between them?

Show the answer

Answer: D. The GCD of 6 and 9 is 3. The 6 ns clock launches at 0, 6, 12; the 9 ns clock captures at 9 and 18. The launch at 6 ns and the capture at 9 ns are 3 ns apart, and no pair is closer.

7.5 Asynchronous clocks and clock groups

Clocks from unrelated sources have no fixed relationship at all, so any timing check between them is meaningless. They are declared as separate clock groups, and the crossings are made safe by design instead.

Suppose one clock is 10 ns and the other is 7.3 ns, from two different crystals. Use the GCD rule anyway and see what happens:

Clocks Edges line up again every Tightest pair
10 ns and 7.3 ns 730 ns 0.100 ns
10 ns and exactly 137 MHz 1000 ns 0.073 ns

The tool would demand that paths between them fit in a few tens of picoseconds. And even that is fiction: two free-running crystals drift, so the edges do not really repeat at all. Given enough time, some edge pair will land as close together as you like.


set_clock_groups -asynchronous -group [get_clocks clk_a] -group [get_clocks clk_b]

This line says: do not time any path between these two groups. The tool stops reporting absurd violations. But it has not made the crossing safe - it has only stopped checking it.

Asynchronous means a CDC job, not a timing job

A signal crossing between unrelated clocks will sometimes change exactly as the capture flip-flop samples it, and metastability follows. That is fixed with a synchroniser and careful crossing design, not with timing constraints. Volume 05 of the Verilog course covers the synchronisers, and State Machines Volume 9.5 shows two state machines talking across two clocks.

Common mistake

Putting two clocks in separate asynchronous groups because paths between them fail. If the clocks are actually related - from the same PLL, at a fixed ratio - this hides real violations. Clock groups are for clocks that truly have no relationship.

Quick check

What does set_clock_groups -asynchronous do to a path between the two groups?

Show the answer

Answer: B. The constraint only removes the timing check between the groups. The crossing still has to be made safe in the design, with synchronisers.

7.6 Clock muxes and clock-gating checks

Clocks that pass through logic need special care. A clock multiplexer carries two clocks that never run at once, so they must not be timed against each other. A clock gate needs its own clock-gating check on its enable.

Clock multiplexers

A clock mux lets the chip run the same logic from clkA (10 ns) or clkB (8 ns). Downstream, the tool sees both clocks arriving at every flip-flop. Left alone, it times paths launched by clkA and captured by clkB - which cannot happen, since only one is ever selected. It would find the 2 ns GCD from sub-module 7.4 and report violations that do not exist.


set_clock_groups -physically_exclusive -group [get_clocks clkA] -group [get_clocks clkB]

Clock gating

Clock gating saves power by stopping the clock to idle flip-flops. The simplest gate is an AND gate: gated clock = clk AND enable. But the enable must not change while the clock is high, or the output pulse is chopped - a glitch on a clock, which is a disaster.

An AND-gate clock gate, with an enable that changes while the clock is high and one that changes while it is low 0 1 2 3 clk en_bad gclk_bad en_ok gclk_ok
Figure 7.4 - Top pair: an enable launched by a rising-edge flip-flop changes while the clock is high, so the gated clock gets chopped pulses. Bottom pair: an enable that only changes while the clock is low gives clean, whole pulses.

So the tool checks the enable against the clock at the gate, like a tiny setup and hold check:

  1. Gating setup: the enable must be settled before the clock rises.
  2. Gating hold: for an AND gate, the enable must not change until the clock has fallen again.

On a 10 ns clock, with the enable coming from 0.20 ns of clock-to-Q and 0.60 ns of logic:

Enable launched by Changes at Gating setup slack Gating hold slack
A rising-edge flip-flop 0.80 ns, while the clock is high 9.20 ns -4.20 ns
A falling-edge flip-flop 5.80 ns, while the clock is low 4.20 ns 0.80 ns

The fix is to change the enable only while the clock is low. Libraries provide an integrated clock gate. It is a latch that holds the enable steady while the clock is high, built together with the AND gate so the timing is right by construction.

Remember

Never build a clock gate from a plain AND gate and an ordinary flip-flop. Use the library's clock-gate cell. The tool checks it automatically, and it cannot glitch.

Quick check

An enable for an AND-gate clock gate is launched by a rising-edge flip-flop. Why does the gating hold check fail?

Show the answer

Answer: C. The flip-flop changes the enable just after the rising edge, while the clock is high. At that moment the AND gate's output follows the enable, so the gated clock pulse is cut short. The enable must wait until the clock falls.

What you learned

Key words from this volume

Every word below has a plain-English entry in the glossary.

Practice

Practice 1

A crossing between 6 ns and 9 ns

A path crosses from a 6 ns clock to a 9 ns clock from the same PLL. It has 0.15 ns of clock-to-Q, 2.60 ns of logic and a 0.08 ns setup time. Does it pass?

Show the solution

The tightest setup check is the GCD of 6 and 9: 3 ns. Slack = 3.00 - 0.15 - 2.60 - 0.08 = 0.17 ns. It passes, but a path of the same length inside either clock domain would have had 6 or 9 ns.

Practice 2

Divided clock, with clock trees

The clk-to-clk_div2 path of sub-module 7.3 is given 0.10 ns of clock tree on both sides. The divider's 0.25 ns stays on the clk_div2 side. What is its slack now?

Show the solution

The 0.10 ns appears on both the launch and the capture side, so it cancels. The divider's 0.25 ns is still extra latency on the capture side. The slack stays at 0.65 ns.

Practice 3

Which constraint?

Match each situation to the right constraint. (a) Two clocks from separate crystals. (b) Two clocks feeding a mux, only one selected at a time. (c) A clock divided by 4 in the design.

Show the solution

(a) set_clock_groups -asynchronous: no timing relationship exists; the crossings need synchronisers.

(b) set_clock_groups -physically_exclusive (or -logically_exclusive): the clocks never coexist on those flip-flops.

(c) create_generated_clock -divide_by 4: the clock is related to its master, and the tool must know how.

Interview corner

Interview question 1

Odd clock ratios

"Data crosses from a 125 MHz clock to a 100 MHz clock, both from one PLL. How much time does the path get?"

Show the solution

"The periods are 8 and 10 ns, and the edges repeat every 40 ns. The tool finds the closest launch and capture pair: a launch at 8 ns captured at 10 ns, so 2 ns - the GCD of the two periods. So a path that looks comfortable in either domain has to fit in 2 ns across the boundary. If it cannot, I would choose a friendlier frequency ratio or add a pipeline stage. Or I would treat the crossing like a CDC, with a handshake or a FIFO."

Interview question 2

Clock gating

"Why is clock gating done with a latch-based integrated clock gate rather than just an AND gate?"

Show the solution

"With a plain AND gate, the enable must not change while the clock is high, or the output pulse is truncated and the downstream flops see a glitch. An enable from a normal rising-edge flop changes just after the rising edge - exactly while the clock is high. The integrated clock gate has a latch that is transparent only while the clock is low, so the enable reaching the AND gate is frozen during the high phase. The tool then performs clock-gating setup and hold checks on the cell, and it is glitch-free by construction."

Volume 08 is about paths that should not be timed the normal way. It covers false paths, multicycle paths, max and min delays, and constant signals - and the damage each one does when it is wrong.