Optimize CPU and IRQ Affinity on Linux Servers

Overview: Why CPU IRQ affinity matters on Linux

On Linux, interrupt handling and CPU scheduling shape network and disk I O performance. By assigning IRQs to specific CPUs and isolating CPUs for application tasks, administrators reduce contention, improve cache locality, and lower latency under high packet or I O load.

This guide focuses on practical steps for system administrators, covering isolcpus, SMT considerations, IRQ affinity via proc, irqbalance behavior, RPS and XPS for network stacks, and how to validate gains with iperf and perf. Commands are given in typical production friendly forms.

Baseline benchmarking and tools

Before changing affinity, collect baseline metrics so you can measure impact. Use iperf3 for network throughput, fio for storage throughput, and perf to capture CPU cycles and IRQ distribution. Record CPU usage, softirq counters, and packet drops.

Useful tools and quick commands:

  • iperf3 for network throughput
  • fio for I O throughput
  • perf top and perf record for hotspots
  • cat /proc/interrupts to inspect IRQ counts

CPU isolation and SMT considerations

Isolate CPUs to reserve cores for latency sensitive workloads. Use the isolcpus kernel parameter at boot, for example isolcpus=2 to isolate logical CPU 2. Combine isolation with cgroups or systemd CPU affinity to ensure system tasks do not migrate onto reserved CPUs.

See also  Linux kpatch Live Kernel Patching: Deploy and Automate

Disable SMT if you observe cross sibling interference on heavy network or storage workloads. SMT off can reduce jitter for some applications, monitor performance with SMT both on and off to decide.

Setting IRQ affinity via proc

Inspect /proc/interrupts to find the IRQ number used by a device. Then write a CPU mask into the IRQ smp affinity file. The mask is a hexadecimal bit mask where bit 0 is CPU0, bit 1 is CPU1, and so on.

# view interrupts
cat /proc/interrupts
# assign IRQ 45 to CPU1 using mask 0x2
echo 2 > /proc/irq/45/smp_affinity

Use a small script to map multiple IRQs to a set of CPUs to balance load across reserved cores without manual repetition.

Managing irqbalance and persistent settings

irqbalance can interfere with manual affinity changes. To test manual settings, stop irqbalance first, then set affinity. If manual tuning helps, either disable irqbalance or pin it to non reserved CPUs via systemd.

# stop irqbalance while tuning
systemctl stop irqbalance
# to prevent restart during testing
systemctl mask irqbalance

For production persistence put affinity changes in a boot script or a udev rule that runs after device initialization, so settings survive reboot.

CPU IRQ affinity Linux

RPS and XPS for network CPU steering

RPS and XPS push packet processing into selected CPUs for better locality. You will write CPU masks into rps_cpus and xps_cpus files under the device queue directories. Use a glob pattern to update all queues at once.

# set RPS mask value 2 for all rx queues of eth0
for f in /sys/class/net/eth0/queues/*/rps_cpus; do
  echo 2 > "$f"
done

Choose masks that map packet processing to isolated CPUs. Test different masks while measuring packet per second and CPU cycles to find the best mapping.

See also  Diagnose and Speed Linux Boot: systemd Unit Tuning

Using taskset and process affinity

Pin critical daemons or user processes to isolated CPUs with taskset using a mask. Avoid allowing background tasks onto reserved cores, so application threads and IRQ handling do not contend with the scheduler.

# pin pid 1234 to CPU1 using mask 0x2
taskset 0x2 1234
# launch a process pinned
taskset 0x2 /usr/bin/myserver

Combine process pinning with IRQ affinity so that the device interrupts and the application both run on nearby CPUs for reduced cache thrash.

Benchmark methodology and expected metrics

Use repeated runs with controlled variables to compare before and after. For network tests run iperf3 for a fixed time and record throughput, retransmits, and CPU usage. For storage run fio with a realistic job file and measure I O per second and latency percentiles.

Key metrics to track: throughput in Gbits or Mbits, I O per second, 99th percentile latency, CPU utilization per core, and /proc/interrupts distribution. Good tuning typically reduces latency and evens IRQ distribution across target CPUs, increasing usable throughput.

Tuning checklist and rollback plan

Follow a controlled rollout. Make one change at a time, measure, then proceed. Keep a rollback plan so you can restore defaults quickly if performance degrades or stability issues appear.

  • Record baseline metrics
  • Stop irqbalance while tuning
  • Set isolcpus and reboot if needed
  • Assign IRQ smp affinity masks and set RPS and XPS via glob updates
  • Pin important processes with taskset
  • Reenable irqbalance only if you configured it to respect your affinity plan

FAQ

This FAQ addresses common practical questions during tuning and deployment.

  • Q: Will disabling irqbalance always help performance?
    A: No, irqbalance is useful on many systems. Disable it only for testing or when you perform manual pinning that gives measurable gains. Reenable if manual tuning is not beneficial.
  • Q: How do I calculate CPU masks?
    A: CPU masks are hex bit masks with bit n representing CPU n. For CPU1 the mask is 0x2, for CPU0 and CPU1 it is 0x3, and so on. Convert binary to hex for larger CPU sets.
  • Q: Is isolcpus enough to prevent scheduler interference?
    A: isolcpus helps, but also use cgroups or systemd CPU affinity and avoid running non critical daemons on isolated CPUs for best results.
  • Q: How do I persist affinity changes across reboot?
    A: Use a boot script or udev rules to write masks at device initialization, or configure irqbalance with a policy that matches your topology if you want a managed approach.
See also  Linux SSH Hardening Guide: Keys, Rate Limits, Audit

Conclusion

Optimizing CPU and IRQ affinity on Linux is a straightforward way to improve network and storage throughput when workloads saturate a server. The core steps are baseline measurement, isolating CPUs, assigning IRQ affinity, steering network processing with RPS and XPS, and pinning important processes. Each change should be tested and measured, and changes need a persistence plan for reboot survival.

Keep in mind that every workload and hardware topology is different. What helps a dual socket NUMA server may not help a single socket cloud instance. Use the checklist in this article to iterate safely, and document your affinity maps so team members can maintain the configuration. With careful measurement and incremental tuning, you can reduce latency, distribute interrupt load sensibly, and get measurable throughput gains on production Linux servers.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top