Input and Output Timing
Half of every input and output path lies outside your design, on a board and inside another chip the timing tool cannot see. So you describe that half with numbers taken from the other chip's datasheet: an input delay and an output delay. This volume turns datasheets into constraints, and then shows why fast interfaces stop sharing one clock and start sending their own.
- How to build set_input_delay from a datasheet, for setup and for hold
- How to build set_output_delay, including a negative minimum
- What a virtual clock is, and when you need one
- Why a system-synchronous interface runs out of speed, and how source-synchronous fixes it
- How DDR halves the time per bit, and how its margins are worked out
10.1 Input delay
An input delay describes the half of an input path that lies outside your design. It is built from the sending chip's datasheet and the board. It is how long after the clock edge the data leaves that chip, plus how long it takes to cross the board.
A timing tool knows everything inside your design and nothing outside it. For an input path, the data was launched by a flip-flop in another chip, on the same board clock, and spent time getting to your pin. You tell the tool how much.
From the datasheet to the constraint
The sending chip's datasheet gives its clock-to-out: at most 4.0 ns, at least 1.0 ns. The board trace takes between 0.5 and 0.8 ns.
| Value | Built from | Used for |
|---|---|---|
| set_input_delay -max 4.80 | 4.0 slowest clock-to-out + 0.8 longest trace | Setup |
| set_input_delay -min 1.50 | 1.0 fastest clock-to-out + 0.5 shortest trace | Hold |
set_input_delay -clock clk -max 4.80 [get_ports in_data]
set_input_delay -clock clk -min 1.50 [get_ports in_data]
Timing the path
Inside our design, the pad and logic take 1.20 ns (0.40 fastest). The clock tree reaches the capture flip-flop after 0.60 ns. Setup is 0.10 ns and hold is 0.05 ns, on a 10 ns clock.
| Check | Arrival | Required | Slack |
|---|---|---|---|
| Setup | 4.80 + 1.20 = 6.00 ns | 10 + 0.60 - 0.10 = 10.50 ns | 4.50 ns |
| Hold | 1.50 + 0.40 = 1.90 ns | 0 + 0.60 + 0.05 = 0.65 ns | 1.25 ns |
So the most logic this input can go through inside our chip, on setup, is 10.00 + 0.60 - 0.10 - 4.80 = 5.70 ns.
Giving only one input delay. A single value is used for both setup and hold, so either the setup check is too kind or the hold check is too harsh. Always give -max and -min, each from its own datasheet figures.
A sending chip's clock-to-out is at most 3.2 ns and at least 2.5 ns. The trace takes 0.3 to 0.6 ns. What is set_input_delay -max?
Show the answer
Answer: B. The maximum uses the slowest numbers: 3.2 + 0.6 = 3.8 ns. The minimum would be 2.5 + 0.3 = 2.8 ns, used for the hold check.
10.2 Output delay
An output delay describes what the receiving chip needs after your pin. Its maximum is the trace plus the receiver's setup time. Its minimum is the trace less the receiver's hold time - and that is often negative.
Now the other direction. Our flip-flop launches data out of a pin, across the board, into another chip's flip-flop. That chip's datasheet gives its setup time, 2.0 ns, and hold time, 0.5 ns. The trace takes 0.4 to 0.7 ns.
| Value | Built from | Used for |
|---|---|---|
| set_output_delay -max 2.70 | 0.7 longest trace + 2.0 their setup | Setup |
| set_output_delay -min -0.10 | 0.4 shortest trace - 0.5 their hold | Hold |
The minimum is negative, and that is correct. Their hold time is longer than the shortest trace, so our data must stay put for a little after the clock edge: 0.10 ns.
Timing the path
Our clock tree reaches the launch flip-flop after 0.60 ns. Clock-to-Q is 0.30 ns (0.20 fastest), and the logic and output pad take 2.10 ns (0.90 fastest).
| Check | Arrival | Required | Slack |
|---|---|---|---|
| Setup | 0.60 + 0.30 + 2.10 = 3.00 ns | 10 - 2.70 = 7.30 ns | 4.30 ns |
| Hold | 0.60 + 0.20 + 0.90 = 1.70 ns | 0 + 0.10 = 0.10 ns | 1.60 ns |
Notice that our own clock tree, 0.60 ns, is on the arrival side only. Volume 06 showed why: the other chip does not share our tree, so our clock latency makes our outputs late.
Input delay: how much of the cycle the outside world has already used. Output delay: how much of the cycle the outside world still needs. Your design gets whatever is left in between.
Leaving a negative minimum output delay out, because a negative delay looks like an error. Without it, the hold check at the output uses zero, and the tool will not notice that the receiver needs the data held after the edge.
A receiving chip needs 1.5 ns of setup and 0.8 ns of hold. The trace takes 0.5 to 0.9 ns. What are set_output_delay -max and -min?
Show the answer
Answer: D. Maximum: the longest trace plus their setup, 0.9 + 1.5 = 2.4 ns. Minimum: the shortest trace less their hold, 0.5 - 0.8 = -0.3 ns.
10.3 Virtual clocks
A virtual clock is a clock with no pin in the design. It stands for the other chip's clock, so input and output delays can be measured from the edge that really launched or captures the data.
So far the other chip ran on our clock, edge for edge. Suppose instead its clock rises 2 ns after ours, because it comes from a different buffer on the board. Our design has no pin carrying that clock - so we declare a virtual one.
create_clock -name vclk -period 10 -waveform {2 7}
set_input_delay -clock vclk -max 4.80 [get_ports in_data]
set_input_delay -clock vclk -min 1.50 [get_ports in_data]
Now the tool pairs the edges properly, using the rule from Volume 07:
| Launch edge (vclk) | Capture edge (clk) | Slack | |
|---|---|---|---|
| Setup | 2 ns | 10 ns | 2.50 ns (was 4.50 ns) |
| Hold | 2 ns | 0 ns | 3.25 ns (was 1.25 ns) |
The input really does have 2 ns less for setup, and 2 ns more margin for hold. Tying the delay to our own clock would have hidden both.
Use a virtual clock whenever the chip at the other end of an interface runs on a clock that is not exactly one of yours. It may be shifted, at a different frequency, or simply not present in your design.
Why is a virtual clock "virtual"?
Show the answer
Answer: A. A virtual clock is created with no source pin. Nothing in the design is clocked by it; it exists so that input and output delays can refer to the outside chip's real clock edges.
10.4 System-synchronous versus source-synchronous
When both chips share one board clock, the data's whole trip must fit in a cycle, and speed runs out quickly. Sending the clock along with the data makes the trip cancel - which is how every fast interface works.
System-synchronous: one clock for everyone
A system-synchronous interface is the one in 10.1: a single board clock drives both chips. Everything must fit inside one clock period:
- the sender's clock-to-out: 4.00 ns
- the board trace: 0.80 ns
- the skew between the board clock at the two chips: 0.50 ns
- our input path and setup time: 1.30 ns
That adds up to 6.60 ns, so this interface can run no faster than 151.5 MHz. And every term is fixed by the board and the other chip. There is little you can do inside your design.
Source-synchronous: send the clock with the data
In a source-synchronous interface, the sending chip forwards its own clock on a pin next to the data. The clock and the data leave together and travel matched traces.
So the clock-to-out and the trace delay affect both equally, and cancel. What is left is only the skew between data and clock - the small differences between the two traces and the two output buffers.
Take a skew of up to 0.30 ns either way, and a receiver setup and hold of 0.20 ns each. Then one bit needs 0.30 + 0.30 + 0.20 + 0.20 = 1.00 ns. That is 1000 Mb/s per pin, against about 151 for the system-synchronous version.
The faster an interface, the more certain it is to be source-synchronous. PCI Express goes further and hides the clock inside the data itself, but the reason is the same: never make the data race a clock that took a different road.
Why can a source-synchronous interface run so much faster than a system-synchronous one?
Show the answer
Answer: C. The forwarded clock suffers the same delays as the data. Only the mismatch between them - the skew - eats into the bit time, and that is far smaller than the whole trip.
10.5 DDR interfaces
A DDR interface sends data on both clock edges, so each bit gets half a period. The receiver captures each bit with a strobe whose edges sit in the middle of the bit, leaving equal margin on both sides.
With the strobe centred, each side of the bit gets half of it. From that half come the skew and the receiver's setup time (or hold time, on the other side). With a skew of 0.30 ns and a setup of 0.20 ns:
| Clock | Period | Bit time | Setup margin |
|---|---|---|---|
| 200 MHz | 5.000 ns | 2.500 ns | 0.750 ns |
| 400 MHz | 2.500 ns | 1.250 ns | 0.125 ns |
| 500 MHz | 2.000 ns | 1.000 ns | 0.000 ns |
| 533 MHz | 1.876 ns | 0.938 ns | -0.031 ns |
| 667 MHz | 1.499 ns | 0.750 ns | -0.125 ns |
At 500 MHz the margin reaches zero: that is the same 1000 Mb/s per pin the source-synchronous sum in 10.4 found. Beyond it, the skew must shrink. Real DDR memories train their strobe delays at power-on, bit by bit, for exactly this reason.
Constraining both edges
DDR inputs need input delays for the rising and the falling edge. In SDC that is done with -clock_fall and -add_delay, so the second pair of numbers is added rather than replacing the first.
Here dqs is the strobe as it arrives at the pins, lined up with the data edges. So the data may change up to 0.30 ns either side of each strobe edge. The receiver then delays the strobe by a quarter period inside the chip, which centres it as in the figure.
set_input_delay -clock dqs -max 0.30 [get_ports dq*]
set_input_delay -clock dqs -min -0.30 [get_ports dq*]
set_input_delay -clock dqs -max 0.30 [get_ports dq*] -clock_fall -add_delay
set_input_delay -clock dqs -min -0.30 [get_ports dq*] -clock_fall -add_delay
Forgetting -add_delay on the falling-edge line. Without it, the second set_input_delay replaces the first, and the rising-edge data is never checked.
A 250 MHz DDR bus has 0.30 ns of skew each way and a 0.20 ns receiver setup time. What is its setup margin?
Show the answer
Answer: B. The period is 4 ns, so each bit is 2 ns and half a bit is 1 ns. Take off the skew and the setup time: 1.000 - 0.300 - 0.200 = 0.500 ns.
What you learned
- Input delay = the sender's clock-to-out plus the board trace: -max for setup, -min for hold.
- Output delay = the trace plus the receiver's setup (-max), or the trace less its hold (-min).
- A negative minimum output delay is normal, and must not be left out.
- Our own clock tree latency helps inputs and hurts outputs.
- A virtual clock describes an outside chip's clock so its edges are paired correctly.
- System-synchronous interfaces run out of speed; source-synchronous ones cancel the trace delay.
- DDR halves the time per bit, and centres a strobe in each bit to share the margin.
Key words from this volume
Every word below has a plain-English entry in the glossary.
- Input delay
- Clock-to-out
- Output delay
- Virtual clock
- System-synchronous
- Source-synchronous
- DDR (double data rate)
- Strobe
Practice
An input on an 8 ns clock
An input has set_input_delay -max 5.6. Inside, it goes through 2.1 ns of logic to a flip-flop with a 0.1 ns setup time. The clock tree to that flip-flop is 0.5 ns. The clock is 8 ns. What is the slack?
Show the solution
Arrival: 5.6 + 2.1 = 7.7 ns. Required: 8.0 + 0.5 - 0.1 = 8.4 ns. Slack = 0.70 ns.
The clock tree helped by 0.5 ns. Without it, the path would only just pass, with 0.2 ns.
Read the datasheet
A receiving chip's datasheet gives a setup time of 1.5 ns and a hold time of 0.8 ns. The traces take 0.5 to 0.9 ns. Write both output delay constraints, and say which one catches a design that changes its output too soon.
Show the solution
set_output_delay -max 2.4 and set_output_delay -min -0.3.
The -min value catches outputs that change too soon. It makes the tool check that our data stays steady until 0.3 ns after the clock edge. Then the receiver's 0.8 ns hold is met even over the shortest trace.
Interview corner
Where does input delay come from?
"How would you work out set_input_delay for an interface from another chip?"
Show the solution
"From the other chip's datasheet and the board. The maximum is its maximum clock-to-output plus the longest board trace, and it constrains setup. The minimum is its minimum clock-to-output plus the shortest trace, and it constrains hold. The other chip may be clocked by something that is not one of my clocks, at a different phase or frequency. Then I would reference the delays to a virtual clock describing its clock. And I would make sure board clock skew between the two chips is in the numbers somewhere."
Why source-synchronous?
"Why do DDR memories send a strobe with the data instead of using the system clock?"
Show the solution
"Because at those speeds the data's trip across the board is a large part of the bit time, and it varies. If the memory had to be sampled with the system clock, the whole clock-to-out and trace delay would have to fit in a fraction of a nanosecond. Sending a strobe with each byte means the strobe travels the same way as the data, so those delays cancel, and only the skew within the group matters. The controller centres the strobe in each bit - often by training its delay at start-up - to split the remaining margin between setup and hold."
Volume 11 turns to variation. No two chips, and no two gates on one chip, have quite the same delay - so how does a timing tool make sure a design works on all of them?