Share on LinkedInBack to deep dives

Network systems

TCP Engine Mechanics

TCP turns an unreliable packet network into an ordered byte stream. It does that with several cooperating control loops, not one magic reliability switch.

Bytes→Segments→Sequence space→Windows→ACKs→Recovery

A useful mental model: storage plus feedback

The sending application writes bytes into a socket. TCP assigns those bytes positions in a sequence space, keeps unacknowledged data available for possible retransmission, and sends only as much as both the receiver and the network appear able to handle. The receiver reorders what arrives, acknowledges contiguous progress, and delivers a clean stream to its application.

1Sequence numbersIdentify byte positions, not packet IDs.
2ACK feedbackMoves the trusted edge of delivered data.
3Two windowsProtect the receiver and the network independently.
4Timers and evidenceRepair loss without assuming every delay is loss.

We will follow one transfer from an application write through packetization, flight, acknowledgment, and recovery. Each numbered playground appears directly after the theory it tests.

Part 1 / Packaging

TCP carries a byte stream; the network carries packets

An application does not hand TCP a permanent set of packets. It writes an ordered stream of bytes. TCP may group those bytes into segments according to current limits and may packetize them differently when retransmitting. A TCP header contributes ports, sequence and acknowledgment numbers, flags, a receive window, and a checksum. IP adds routing information, while the link layer wraps the IP packet for the next hop.

MTU and MSS describe different boundaries. The MTU is the largest IP packet a particular link can carry without IP fragmentation. MSS is the largest TCP payload an endpoint offers to receive in one segment. With a 1500-byte MTU and ordinary 20-byte IPv4 plus 20-byte TCP headers, the familiar MSS is 1460 bytes. Options or IPv6 change that arithmetic.

Use Playground 1 to enlarge the application write and change the MTU. Observe that the byte stream remains one logical stream even as the segment boundaries move.

Playground 1

Encapsulation and Segmentation

Change the application write and link MTU. TCP divides the byte stream into MSS-sized payloads; IP and Ethernet then add their own envelopes.

Application4000 B byte stream
TCP3 segments, MSS 1460 B
IP20 B IPv4 header per segment
LinkEach IP packet stays at or below 1500 B
SEQ +01460 B+ 40 B headers
SEQ +14601460 B+ 40 B headers
SEQ +29201080 B+ 40 B headers

The final segment may be smaller. MSS is a TCP payload ceiling, while MTU constrains the complete IP packet on one hop.

Part 2 / Sequence space

The sender and receiver each track a moving interval

TCP sequence numbers count bytes modulo 232. At the sender, SND.UNA is the oldest unacknowledged byte andSND.NXT is the next byte to send. At the receiver,RCV.NXT is the next contiguous byte expected. An ACK number means, “I have everything before this position.” SYN and FIN each consume one sequence number even though they are not payload.

ACKedin flightallowed nextnot allowed yet

SND.UNASND.NXTSND.UNA + window

Data in flight is data sent but not cumulatively acknowledged. The sender must retain enough state to recover it. Once an ACK advancesSND.UNA, acknowledged ranges can leave the retransmission accounting and new data can enter the window.

Part 3 / Speed limits

rwnd protects the receiver; cwnd protects the path

The advertised receive window, rwnd, says how many more bytes the receiving TCP is prepared to accept. It reflects receive buffering and application consumption. The congestion window,cwnd, is local sender state inferred from network feedback. It is not transmitted as a TCP header field.

The sender may keep no more than roughlymin(cwnd, rwnd) bytes outstanding. That is a flight limit, not the full throughput formula. A path with bandwidthB and round-trip time RTT can hold aboutB × RTT bits. This bandwidth-delay product is the amount of in-flight data required to keep the path continuously busy.

Use Playground 2 to create a long-delay path and see why a window that feels large on a LAN can starve a fast WAN.

Playground 2

The Flight Window and the Bandwidth-Delay Product

Build a path, then compare how many bytes the network can hold with how many bytes TCP is allowed to keep unacknowledged.

network pipe: 500.0 KB
Usable flight window256.0 KB
Estimated ceiling51.2 Mbps
Path utilization51%
Current limitercongestion window

min(cwnd, rwnd) limits bytes in flight. Throughput is then bounded by roughly flight / RTT, the path capacity, and the application's ability to provide or consume data.

Part 4 / Congestion control

ACKs are both receipts and a clock

Returning ACKs prove that earlier data left the network and reached the receiver. That feedback releases room for more data. During classic slow start, ACKed data grows cwnd rapidly, approximately doubling it each RTT. Above the slow-start threshold, congestion avoidance grows more cautiously.

A loss-based algorithm treats loss or ECN as evidence that the path was pushed too far. Reno uses additive increase and multiplicative decrease. CUBIC follows a time-based cubic growth function and is the common Linux default. BBR instead models bottleneck bandwidth and propagation RTT and uses pacing; it is not simply “CUBIC without reacting to loss.” Every algorithm still has to coexist with path capacity, queues, receiver limits, and competing flows.

Playground 3 intentionally uses a small Reno-style teaching model. Move the threshold and loss point to see why slow start and congestion avoidance have visibly different slopes.

Playground 3

ACK Clock and Congestion Window

Step through a teaching model of Reno-style growth. ACKs release new data; a loss signal cuts the sending window and changes the slope.

21
42
83
164
175
186
197
208
109
1110
1211
1312
ACK-driven growthwindow after loss

This isolates the classic idea: exponential growth belowssthresh, additive growth above it, and multiplicative decrease after congestion. Production CUBIC, BBR, pacing, PRR, and modern loss detection add more machinery.

Part 5 / Time

The retransmission timeout learns the path

A sender cannot label a packet lost merely because it has not yet arrived. It maintains a smoothed RTT (SRTT) and RTT variation (RTTVAR), then places the retransmission timeout beyond the expected delay. The classic estimator isRTO = SRTT + 4 × RTTVAR, subject to implementation and standards bounds. RFC 6298 recommends an initial and minimum RTO of one second; production stacks may apply more recent loss detection and implementation-specific timer floors.

TCP conceptually times the oldest outstanding data rather than needing one independent protocol timer for every segment. After a timeout, the sender retransmits and exponentially backs off the RTO. Karn's algorithm avoids taking ambiguous RTT samples from retransmitted data.

In Playground 4, add jitter and one delay spike. The visualization uses a 200 ms teaching floor so changes remain visible; focus on how variation expands the safety margin.

Playground 4

One Retransmission Timer, Continuously Recalibrated

Disturb the RTT samples. The timeout follows both the smoothed RTT and its variation, so a noisy path needs a larger safety margin.

100RTT 1
110RTT 2
95RTT 3
115RTT 4
100RTT 5
oldest unacknowledged byte
RTO 200 ms
SRTT102 ms
RTT variation21 ms
Safety margin98 ms

Part 6 / Loss recovery

Duplicate ACKs reveal a gap before the timer expires

Suppose range 2 is lost but ranges 3, 4, and 5 arrive. The receiver continues acknowledging the next missing range. Three duplicate ACKs are the classic trigger for fast retransmit, allowing repair without waiting for RTO. Modern stacks may also use algorithms such as RACK, which reasons about packet timing and reordering.

A cumulative ACK alone says where contiguous delivery stops. The SACK option adds blocks describing later ranges already received. This scoreboard lets the sender target multiple holes instead of rediscovering each one over successive RTTs. SACK is negotiated in the handshake; it does not change the cumulative meaning of the ACK field.

Use Playground 5 to remove one or several segments, then disable SACK. Compare the evidence available to the sender.

Playground 5

Fast Retransmit, Cumulative ACKs, and SACK

Click packets to lose or restore them. The cumulative ACK identifies the first missing byte range; SACK reports later ranges already held.

Sender123456
Receiver1gap3456
Cumulative ACKACK 2next contiguous segment wanted
SACK option3-6later ranges safely buffered
Retransmit2scoreboard can target gaps

Part 7 / Path MTU

When a perfectly valid segment is too large for one hop

The path MTU is the smallest MTU along a route. With IPv4's Don't Fragment bit or with IPv6, an oversized packet cannot simply pass the narrow hop. Traditional PMTUD depends on ICMP feedback reaching the sender. If a firewall discards that feedback, large packets may disappear while small handshake packets still work: a classic PMTU blackhole.

MSS clamping rewrites the MSS offered in SYN packets so endpoints packetize conservatively from the beginning. It is a pragmatic network workaround, not proof that ICMP should be blocked. Packetization-layer PMTUD can instead probe sizes end to end without depending solely on ICMP.

Toggle ICMP and clamping in Playground 6. Watch why the connection may establish successfully yet stall only when it begins sending full-sized data.

Playground 6

Path MTU Blackhole

Put a smaller hop in the path. Then decide whether its ICMP feedback reaches the sender or a router clamps the MSS during the handshake.

Sender1500 B packet
×
MTU 1400packet too large
×
Receiverstill waiting

The packet is dropped and feedback is hidden: repeated retransmission can look like a hanging connection.

Part 8 / Queues

A full queue can preserve throughput while destroying latency

A router serializes packets onto a finite-rate link. If arrivals exceed departures, a queue grows. Deep unmanaged buffers can hide congestion by delaying packets for hundreds of milliseconds before finally dropping them. Bulk throughput may remain high while DNS, SSH, calls, and games become painfully slow.

Active Queue Management (AQM), commonly paired with fair queueing, marks or drops selected packets before the queue becomes enormous. The aim is not an empty router at every instant; it is enough buffering to absorb useful bursts without turning persistent load into persistent delay.

Use Playground 7 to overload the bottleneck. Compare a large FIFO with an idealized low-delay AQM.

Playground 7

Bufferbloat and Active Queue Management

Offer more traffic than a 100 Mbps bottleneck can drain. A large FIFO preserves packets by making every flow wait; AQM signals congestion before delay becomes enormous.

120Mbps in
320 ms queued
100Mbps out
Base RTT20 ms
Observed RTT340 ms
Interactive feelpainful

Part 9 / Synthesis

One transfer, many cooperating mechanisms

The application supplies bytes. Packetization respects path and peer limits. Sequence numbers preserve order. rwnd protects receiver memory, while cwnd and pacing regulate pressure on the path. ACKs advance delivery, measure time, and clock new work. SACK and loss detection repair gaps. Router queues absorb bursts but become harmful when they conceal persistent overload.

Run Playground 8 slowly. At each step, identify which state belongs to the sender, which is advertised by the receiver, and which exists only inside the network.

Playground 8

Run the Whole TCP Engine

Advance one event at a time. Watch storage, reliability, flow control, congestion control, and the router queue cooperate across one flight.

Sendersend + retransmission buffers
Routerbottleneck queue
Receiverreassembly buffer
Step 1 of 6Application writes

The socket send buffer receives an ordered byte stream.

Keep these distinctions

MTU is not MSS.MTU constrains an IP packet on a link; MSS constrains TCP payload accepted by an endpoint.
rwnd is not cwnd.The receiver advertises rwnd; the sender maintains cwnd from congestion-control state.
Window is not throughput.A flight window interacts with RTT, path capacity, pacing, loss, and application behavior.
Delay is not necessarily loss.TCP combines acknowledgments, timers, reordering evidence, and optional SACK information.
A full pipe is not a full queue.BDP represents useful in-flight data; queue buildup is extra waiting at a bottleneck.

Primary references