Introduction and goals for Linux TCP tuning
This guide targets Linux administrators and network engineers who need practical TCP tuning to maximize throughput and minimize latency on servers. It focuses on kernel and NIC level adjustments, plus measurement methods so you can tune with data rather than guesswork.
We cover baseline benchmarking, essential sysctl parameters, socket buffers, IRQ affinity and CPU assignment, NIC offloads and ethtool tweaks, RPS and XPS, plus iperf based testing and monitoring. Follow these steps in a lab before rolling changes to production.
Baseline benchmarking and metrics
Start with a baseline so you can measure impact. Key metrics are throughput in Gbps or Mbps, round trip time and packet loss. Use iperf3 for TCP throughput, single threaded and parallel stream tests, and ping or hping3 for latency and loss profiles.
Collect system metrics while testing: CPU usage, context switches, softirqs, NIC queue drops and ring buffer statistics. Tools to use include sar, vmstat, ss, ethtool -S and top or htop on busy systems.
Kernel sysctl settings that matter
Tune the kernel network stack by adjusting memory limits, backlog values and TCP behavior. Typical settings to evaluate are net.core.rmem_max, net.core.wmem_max, net.core.netdev_max_backlog and several tcp memory controls. Apply changes temporarily with sysctl or persist them in /etc/sysctl.conf.
Recommended parameters to test include increasing receive and send buffers, enabling larger backlog, and selecting an appropriate congestion control algorithm such as cubic or bbr depending on your workload and kernel version. Record each change so you can compare results.
Socket buffer tuning and application knobs
Sockets inherit kernel limits, so increase defaults when your application uses many concurrent connections or high throughput. Adjust net.ipv4.tcp_rmem and net.ipv4.tcp_wmem to raise min, default and max values. Ensure applications are configured to use auto tunings where available.
For servers handling large transfers, increase accept queues and tune application level timeouts. Use ss to inspect socket buffer usage and tune per process if needed. When possible, enable TCP window scaling at the application or socket layer to make use of larger buffers.
NIC offloads and ethtool tweaks
NIC offloads can reduce CPU cost but sometimes increase latency or cause poor behavior with small packets. Test enabling and disabling GRO, GSO and TSO to see which configuration matches your latency and throughput goals. Use ethtool to change offload settings and to inspect ring sizes.
Also review interrupt moderation and coalescing settings with ethtool -C, and experiment with changing RX and TX ring sizes. For some workloads, increasing rings reduces packet drops, for others reducing coalescing can improve latency.

IRQ affinity and CPU isolation
Pin NIC interrupts to specific CPUs to avoid cross CPU contention and cache thrash. Use irqbalance for automated distribution or set smp_affinity manually for deterministic placement. Align IRQs with CPU cores that run your network heavy processes.
Consider isolating CPUs with kernel boot parameters for real time like behavior under heavy network load. Avoid moving user processes around while tuning, and validate affinity changes with /proc/interrupts and tools such as perf or vmstat.
RSS, RPS and XPS tuning
Receive side scaling and transmit packet steering distribute packet processing across CPUs. Configure RSS in the NIC if hardware supports it, and tune RPS and XPS via sysfs entries under /sys/class/net/INTERFACE/queues. Distribute flows across cores that are not overloaded.
Balance RPS masks and flow counts to avoid single CPU saturation. For virtualized environments, ensure vCPU count and queue pairs are aligned with physical NIC settings, and test varying the number of channels with ethtool -L to match throughput targets.
Testing with iperf and interpreting results
Use iperf3 with different stream counts to see how throughput scales. Run single stream tests, then parallel streams with -P to simulate multiple flows. Use -O to skip warm up and record throughput steady state. Combine with ping to track latency during the same window.
Interpret results by comparing CPU levels, queue drops and softirq counts between runs. A higher throughput with a big rise in softirqs or context switches points to CPU or interrupt tuning needs rather than socket buffer limits.
Monitoring, rollout strategy and rollback
Create a staged rollout plan, move changes from lab to a canary host and then to wider production. Monitor application latency, retransmits, and NIC error counters after each step. Keep a clear rollback playbook to restore previous sysctl or ethtool settings quickly if an issue appears.
Automate persistent changes with configuration management, and document every tweak along with the observed impact. Continuous monitoring with Prometheus, Grafana, or similar systems helps catch regressions, and alerting thresholds should include retransmit rate and packet drops.
FAQs
Here are common quick answers to practical questions that come up during Linux TCP tuning. These are short practical notes, not exhaustive explanations.
Q: Will the same tuning work for all NICs and kernels? A: No, hardware, driver quality and kernel version matter. Test per platform and driver, adjust IRQ, RSS and offload settings accordingly.
Q: How do I revert bad sysctl or ethtool changes? A: Keep a saved copy of original settings, use sysctl -w to revert values for the running system, and ethtool to restore NIC options. Automate restores via your config management tool.
- Q: Should I disable offloads to reduce latency? A: Some workloads benefit from disabling GRO GSO and TSO, especially small packet, low latency traffic. For bulk throughput, hardware offloads usually improve CPU efficiency.
- Q: Which congestion control should I pick? A: cubic is default and stable. bbr can improve throughput and latency in high bandwidth delay product links, test it under realistic load before enabling.
Conclusion and next steps
Tuning the Linux TCP stack is a balance between throughput, latency and CPU utilization, and it requires data driven iteration. Begin with controlled benchmarks, collect CPU and NIC metrics, then tune memory, backlog and offloads. Adjust IRQ affinity and scaling features to match the core topology of the host, and always validate each change with iperf and system counters.
Document every adjustment, automate persistent changes and stage rollouts to minimize production risk. Build dashboards that surface retransmits, NIC drops and softirq spikes so regressions are visible quickly. With methodical testing and incremental changes you will find the combination of sysctl, socket settings and NIC configuration that matches your performance goals. Keep a rollback process ready and schedule maintenance windows for risky experiments. Over time, create a tuning profile per workload class so you can apply a tested set of parameters to new servers consistently, saving time and reducing variability across your fleet.