Volume 10 Intermediate 5 sub-modules ~20 min read

Input and Output Timing

Half of every input and output path lies outside your design, on a board and inside another chip the timing tool cannot see. So you describe that half with numbers taken from the other chip's datasheet: an input delay and an output delay. This volume turns datasheets into constraints, and then shows why fast interfaces stop sharing one clock and start sending their own.

You will learn
  • How to build set_input_delay from a datasheet, for setup and for hold
  • How to build set_output_delay, including a negative minimum
  • What a virtual clock is, and when you need one
  • Why a system-synchronous interface runs out of speed, and how source-synchronous fixes it
  • How DDR halves the time per bit, and how its margins are worked out
You need

10.1 Input delay

An input delay describes the half of an input path that lies outside your design. It is built from the sending chip's datasheet and the board. It is how long after the clock edge the data leaves that chip, plus how long it takes to cross the board.

A timing tool knows everything inside your design and nothing outside it. For an input path, the data was launched by a flip-flop in another chip, on the same board clock, and spent time getting to your pin. You tell the tool how much.

An input path crossing from a sending chip, over a board trace, into our design sending chip our design FF pad + logic FF board trace one board clock feeds both chips input delay: clock-to-out + trace timed by the tool
Figure 10.1 - The input delay covers everything outside our design: the sending chip's clock-to-out and the board trace. The tool times the rest - our pad, our logic and our capture flip-flop - itself.

From the datasheet to the constraint

The sending chip's datasheet gives its clock-to-out: at most 4.0 ns, at least 1.0 ns. The board trace takes between 0.5 and 0.8 ns.

Value Built from Used for
set_input_delay -max 4.80 4.0 slowest clock-to-out + 0.8 longest trace Setup
set_input_delay -min 1.50 1.0 fastest clock-to-out + 0.5 shortest trace Hold

set_input_delay -clock clk -max 4.80 [get_ports in_data]
set_input_delay -clock clk -min 1.50 [get_ports in_data]

Timing the path

Inside our design, the pad and logic take 1.20 ns (0.40 fastest). The clock tree reaches the capture flip-flop after 0.60 ns. Setup is 0.10 ns and hold is 0.05 ns, on a 10 ns clock.

Check Arrival Required Slack
Setup 4.80 + 1.20 = 6.00 ns 10 + 0.60 - 0.10 = 10.50 ns 4.50 ns
Hold 1.50 + 0.40 = 1.90 ns 0 + 0.60 + 0.05 = 0.65 ns 1.25 ns

So the most logic this input can go through inside our chip, on setup, is 10.00 + 0.60 - 0.10 - 4.80 = 5.70 ns.

Common mistake

Giving only one input delay. A single value is used for both setup and hold, so either the setup check is too kind or the hold check is too harsh. Always give -max and -min, each from its own datasheet figures.

Quick check

A sending chip's clock-to-out is at most 3.2 ns and at least 2.5 ns. The trace takes 0.3 to 0.6 ns. What is set_input_delay -max?

Show the answer

Answer: B. The maximum uses the slowest numbers: 3.2 + 0.6 = 3.8 ns. The minimum would be 2.5 + 0.3 = 2.8 ns, used for the hold check.

10.2 Output delay

An output delay describes what the receiving chip needs after your pin. Its maximum is the trace plus the receiver's setup time. Its minimum is the trace less the receiver's hold time - and that is often negative.

Now the other direction. Our flip-flop launches data out of a pin, across the board, into another chip's flip-flop. That chip's datasheet gives its setup time, 2.0 ns, and hold time, 0.5 ns. The trace takes 0.4 to 0.7 ns.

Value Built from Used for
set_output_delay -max 2.70 0.7 longest trace + 2.0 their setup Setup
set_output_delay -min -0.10 0.4 shortest trace - 0.5 their hold Hold

The minimum is negative, and that is correct. Their hold time is longer than the shortest trace, so our data must stay put for a little after the clock edge: 0.10 ns.

Timing the path

Our clock tree reaches the launch flip-flop after 0.60 ns. Clock-to-Q is 0.30 ns (0.20 fastest), and the logic and output pad take 2.10 ns (0.90 fastest).

Check Arrival Required Slack
Setup 0.60 + 0.30 + 2.10 = 3.00 ns 10 - 2.70 = 7.30 ns 4.30 ns
Hold 0.60 + 0.20 + 0.90 = 1.70 ns 0 + 0.10 = 0.10 ns 1.60 ns

Notice that our own clock tree, 0.60 ns, is on the arrival side only. Volume 06 showed why: the other chip does not share our tree, so our clock latency makes our outputs late.

In plain words

Input delay: how much of the cycle the outside world has already used. Output delay: how much of the cycle the outside world still needs. Your design gets whatever is left in between.

Common mistake

Leaving a negative minimum output delay out, because a negative delay looks like an error. Without it, the hold check at the output uses zero, and the tool will not notice that the receiver needs the data held after the edge.

Quick check

A receiving chip needs 1.5 ns of setup and 0.8 ns of hold. The trace takes 0.5 to 0.9 ns. What are set_output_delay -max and -min?

Show the answer

Answer: D. Maximum: the longest trace plus their setup, 0.9 + 1.5 = 2.4 ns. Minimum: the shortest trace less their hold, 0.5 - 0.8 = -0.3 ns.

10.3 Virtual clocks

A virtual clock is a clock with no pin in the design. It stands for the other chip's clock, so input and output delays can be measured from the edge that really launched or captures the data.

So far the other chip ran on our clock, edge for edge. Suppose instead its clock rises 2 ns after ours, because it comes from a different buffer on the board. Our design has no pin carrying that clock - so we declare a virtual one.


create_clock -name vclk -period 10 -waveform {2 7}
set_input_delay -clock vclk -max 4.80 [get_ports in_data]
set_input_delay -clock vclk -min 1.50 [get_ports in_data]

Now the tool pairs the edges properly, using the rule from Volume 07:

Launch edge (vclk) Capture edge (clk) Slack
Setup 2 ns 10 ns 2.50 ns (was 4.50 ns)
Hold 2 ns 0 ns 3.25 ns (was 1.25 ns)

The input really does have 2 ns less for setup, and 2 ns more margin for hold. Tying the delay to our own clock would have hidden both.

Remember

Use a virtual clock whenever the chip at the other end of an interface runs on a clock that is not exactly one of yours. It may be shifted, at a different frequency, or simply not present in your design.

Quick check

Why is a virtual clock "virtual"?

Show the answer

Answer: A. A virtual clock is created with no source pin. Nothing in the design is clocked by it; it exists so that input and output delays can refer to the outside chip's real clock edges.

10.4 System-synchronous versus source-synchronous

When both chips share one board clock, the data's whole trip must fit in a cycle, and speed runs out quickly. Sending the clock along with the data makes the trip cancel - which is how every fast interface works.

System-synchronous: one clock for everyone

A system-synchronous interface is the one in 10.1: a single board clock drives both chips. Everything must fit inside one clock period:

  1. the sender's clock-to-out: 4.00 ns
  2. the board trace: 0.80 ns
  3. the skew between the board clock at the two chips: 0.50 ns
  4. our input path and setup time: 1.30 ns

That adds up to 6.60 ns, so this interface can run no faster than 151.5 MHz. And every term is fixed by the board and the other chip. There is little you can do inside your design.

Source-synchronous: send the clock with the data

In a source-synchronous interface, the sending chip forwards its own clock on a pin next to the data. The clock and the data leave together and travel matched traces.

So the clock-to-out and the trace delay affect both equally, and cancel. What is left is only the skew between data and clock - the small differences between the two traces and the two output buffers.

Take a skew of up to 0.30 ns either way, and a receiver setup and hold of 0.20 ns each. Then one bit needs 0.30 + 0.30 + 0.20 + 0.20 = 1.00 ns. That is 1000 Mb/s per pin, against about 151 for the system-synchronous version.

Remember

The faster an interface, the more certain it is to be source-synchronous. PCI Express goes further and hides the clock inside the data itself, but the reason is the same: never make the data race a clock that took a different road.

Quick check

Why can a source-synchronous interface run so much faster than a system-synchronous one?

Show the answer

Answer: C. The forwarded clock suffers the same delays as the data. Only the mismatch between them - the skew - eats into the bit time, and that is far smaller than the whole trip.

10.5 DDR interfaces

A DDR interface sends data on both clock edges, so each bit gets half a period. The receiver captures each bit with a strobe whose edges sit in the middle of the bit, leaving equal margin on both sides.

A DDR data bus with two bits per clock period, and a strobe centred in each bit 0 1 2 3 clk dq D0 D1 D2 D3 D4 D5 D6 D7 dqs
Figure 10.2 - Two bits per clock period: one after each rising edge and one after each falling edge. The strobe is the clock shifted by a quarter of a period, so each of its edges lands in the middle of a bit, where the data is steadiest.

With the strobe centred, each side of the bit gets half of it. From that half come the skew and the receiver's setup time (or hold time, on the other side). With a skew of 0.30 ns and a setup of 0.20 ns:

Clock Period Bit time Setup margin
200 MHz 5.000 ns 2.500 ns 0.750 ns
400 MHz 2.500 ns 1.250 ns 0.125 ns
500 MHz 2.000 ns 1.000 ns 0.000 ns
533 MHz 1.876 ns 0.938 ns -0.031 ns
667 MHz 1.499 ns 0.750 ns -0.125 ns

At 500 MHz the margin reaches zero: that is the same 1000 Mb/s per pin the source-synchronous sum in 10.4 found. Beyond it, the skew must shrink. Real DDR memories train their strobe delays at power-on, bit by bit, for exactly this reason.

Constraining both edges

DDR inputs need input delays for the rising and the falling edge. In SDC that is done with -clock_fall and -add_delay, so the second pair of numbers is added rather than replacing the first.

Here dqs is the strobe as it arrives at the pins, lined up with the data edges. So the data may change up to 0.30 ns either side of each strobe edge. The receiver then delays the strobe by a quarter period inside the chip, which centres it as in the figure.


set_input_delay -clock dqs -max 0.30 [get_ports dq*]
set_input_delay -clock dqs -min -0.30 [get_ports dq*]
set_input_delay -clock dqs -max 0.30 [get_ports dq*] -clock_fall -add_delay
set_input_delay -clock dqs -min -0.30 [get_ports dq*] -clock_fall -add_delay
Common mistake

Forgetting -add_delay on the falling-edge line. Without it, the second set_input_delay replaces the first, and the rising-edge data is never checked.

Quick check

A 250 MHz DDR bus has 0.30 ns of skew each way and a 0.20 ns receiver setup time. What is its setup margin?

Show the answer

Answer: B. The period is 4 ns, so each bit is 2 ns and half a bit is 1 ns. Take off the skew and the setup time: 1.000 - 0.300 - 0.200 = 0.500 ns.

What you learned

Key words from this volume

Every word below has a plain-English entry in the glossary.

Practice

Practice 1

An input on an 8 ns clock

An input has set_input_delay -max 5.6. Inside, it goes through 2.1 ns of logic to a flip-flop with a 0.1 ns setup time. The clock tree to that flip-flop is 0.5 ns. The clock is 8 ns. What is the slack?

Show the solution

Arrival: 5.6 + 2.1 = 7.7 ns. Required: 8.0 + 0.5 - 0.1 = 8.4 ns. Slack = 0.70 ns.

The clock tree helped by 0.5 ns. Without it, the path would only just pass, with 0.2 ns.

Practice 2

Read the datasheet

A receiving chip's datasheet gives a setup time of 1.5 ns and a hold time of 0.8 ns. The traces take 0.5 to 0.9 ns. Write both output delay constraints, and say which one catches a design that changes its output too soon.

Show the solution

set_output_delay -max 2.4 and set_output_delay -min -0.3.

The -min value catches outputs that change too soon. It makes the tool check that our data stays steady until 0.3 ns after the clock edge. Then the receiver's 0.8 ns hold is met even over the shortest trace.

Interview corner

Interview question 1

Where does input delay come from?

"How would you work out set_input_delay for an interface from another chip?"

Show the solution

"From the other chip's datasheet and the board. The maximum is its maximum clock-to-output plus the longest board trace, and it constrains setup. The minimum is its minimum clock-to-output plus the shortest trace, and it constrains hold. The other chip may be clocked by something that is not one of my clocks, at a different phase or frequency. Then I would reference the delays to a virtual clock describing its clock. And I would make sure board clock skew between the two chips is in the numbers somewhere."

Interview question 2

Why source-synchronous?

"Why do DDR memories send a strobe with the data instead of using the system clock?"

Show the solution

"Because at those speeds the data's trip across the board is a large part of the bit time, and it varies. If the memory had to be sampled with the system clock, the whole clock-to-out and trace delay would have to fit in a fraction of a nanosecond. Sending a strobe with each byte means the strobe travels the same way as the data, so those delays cancel, and only the skew within the group matters. The controller centres the strobe in each bit - often by training its delay at start-up - to split the remaining margin between setup and hold."

Volume 11 turns to variation. No two chips, and no two gates on one chip, have quite the same delay - so how does a timing tool make sure a design works on all of them?