Goset is a Linux CLI in Go that selects the CPUs a task runs on. It samples interrupt counts per CPU, ranks cores by the noise of both their SMT threads, and starts the task on the quietest ones. Optional layers add a cgroup v2 cpuset, IRQ steering and a fence on the SMT siblings. Every run reports the kernel counters that explain its variance. A companion tool, goset-bench, compares unpinned, taskset and goset runs of any command.
01Motivation
Run-to-run spread of a CPU-bound benchmark has an operating-system component: interrupts handled on the task's core, activity on the SMT sibling, migrations, preemption. Full isolation (isolcpus, nohz_full) is set on the kernel command line and needs a reboot, which rules it out on a server that cannot go down for a benchmark. The local alternative, taskset, pins the task but leaves the choice of CPU to the user, with no view of what each CPU is doing. Goset aims at both: convenient, one command at runtime, and optimized, with the CPU picked from a measurement.
goset -- ./task # one quiet CPU, no root
goset -n 5 -- ./task # five CPUs, one shared mask
sudo goset -n 1 -cgroup -steer -fence -- ./task02How it works
Ranking
A 500 ms sample of /proc/interrupts gives a per-CPU delta, split into steerable interrupts and non-steerable ones (NMI, local timer, rescheduling IPIs). With -steer only the non-steerable count is noise, since the steerable ones get moved away.
The cost of a CPU is its own noise plus the noise of its SMT sibling. Sort order:
- CPUs named by
-include - cost, ascending
- own noise, then own steerable count
- core id, then CPU id
The task takes the first -n entries. The housekeeper, the thread that polls telemetry, takes the next CPU outside the task's cores.
Isolation
| Mode | Mechanism |
|---|---|
| default | sched_setaffinity on the launching thread before exec, restored after. Every thread the task creates inherits the mask. |
-cgroup | cgroup v2 cpuset, task placed at clone with CLONE_INTO_CGROUP, then pinned to the selected CPUs. Needs root. |
-fence | The SMT siblings of the task CPUs join the cpuset and stay idle. A busy sibling can halve the throughput of an FP-bound task. |
The mask alone does not hold against a task that resets its own affinity (OpenMP, MPI runtimes). The cpuset does.
IRQ steering
-steer stops the irqbalance unit for the run and restarts it on exit. It writes /proc/irq/*/smp_affinity_list for every IRQ currently routed to a selected CPU, then restores the previous affinity. The report counts applied, rejected and remaining IRQs. A run with no systemd-managed irqbalance warns and steers anyway.
Telemetry
A housekeeper thread, pinned outside the task's CPUs, polls kernel software counters: interrupts per CPU, thermal throttle count, CPU frequency, context switches, migrations, run-queue delay. Every source is procfs or sysfs, so no read runs on the task's CPU and no IPI is sent.
Goset reads no PMU counters. Cache, branch and ALU events stay with perf. The two layers answer different questions: goset reports what the OS did to the run, perf reports what the software did.
03Trade-offs
- One interrupt sample. The ranking reads 500 ms of
/proc/interruptsat start. It keeps no history of a CPU across runs. - Software counters only. procfs and sysfs need no PMU access and send no IPI. Cache and branch events stay with
perf. - Plain pinning by default.
sched_setaffinityneeds no privilege and covers multi-threaded tasks. Exclusivity needs-cgroupand root.
04Benchmarks
AMD Ryzen 5 7600X (6 cores, SMT on), boost on. 10 interleaved runs per mode, 5 s idle before each run. Goset mode: -n 1 -cgroup -steer -fence. The taskset mode moves to the next CPU every run, so its spread includes core-to-core differences.
Columns are over the 10 per-run values. spread is (max - min) / median.
jitter: one dependent xorshift-multiply chain, ms per iteration, lower is better.
| mode | median | sd | min | max | spread |
|---|---|---|---|---|---|
| baseline | 5.05 | 0.059 | 4.98 | 5.14 | 3.2% |
| taskset | 5.14 | 0.059 | 5.00 | 5.17 | 3.3% |
| goset | 5.00 | 0.042 | 4.96 | 5.10 | 2.8% |
chase: dependent loads over a 192 KiB ring in L2, ns per load, lower is better.
| mode | median | sd | min | max | spread |
|---|---|---|---|---|---|
| baseline | 2.61 | 0.05 | 2.52 | 2.70 | 6.9% |
| taskset | 2.63 | 0.04 | 2.59 | 2.70 | 4.2% |
| goset | 2.54 | 0.02 | 2.51 | 2.56 | 2.0% |
matmul: 32x32 double matrix multiply, L1-resident, MFLOP/s, higher is better.
| mode | median | sd | min | max | spread |
|---|---|---|---|---|---|
| baseline | 17241 | 597 | 15803 | 17689 | 10.9% |
| taskset | 16556 | 736 | 14856 | 17434 | 15.6% |
| goset | 18191 | 336 | 17134 | 18352 | 6.7% |
stream Triad: 20M-element arrays, MB/s, higher is better.
| mode | median | sd | min | max | spread |
|---|---|---|---|---|---|
| baseline | 40826 | 363 | 40182 | 41367 | 2.9% |
| taskset | 40645 | 395 | 40136 | 41415 | 3.1% |
| goset | 41156 | 233 | 40896 | 41799 | 2.2% |
Medians differ by under 4% between modes. The effect is in the spread: sd is lowest under goset on all four.
05Next
The open issues that change measurement quality most:
- Scheduler-domain exclusion (#87). A cpuset confines the task but leaves its CPUs in the general scheduler domain. Writing
isolatedtocpuset.cpus.partitionremoves them from load balancing. The report states the partition state actually achieved. - IPI audit of the telemetry (#64). A per-source test counts the IPIs each probe causes, driver-backed paths included (the cpufreq read, for one). Only metrics proven IPI-free are exposed.
- Controlled setup (#119, #120).
-stable-envexecs the task with a fixed environment block and-no-aslrfixes the address layout, the two setup variables Mytkowicz et al. found to change measured results. - Reproducibility manifest (#88). Kernel command line, microcode, governor, SMT and ASLR state, and the isolation applied, exported with each result.
06Further reading
- Producing Wrong Data Without Doing Anything Obviously Wrong!: Mytkowicz, Diwan, Hauswirth, Sweeney, ASPLOS 2009. Setup bias from environment size and link order.
- Stabilizer: Statistically Sound Performance Evaluation: Curtsinger, Berger, ASPLOS 2013. Randomizing layout to separate real effects from noise.
- Rigorous Benchmarking in Reasonable Time: Kalibera, Jones, ISMM 2013. Environmental accounting and confidence intervals for benchmark results.
- The Case of the Missing Supercomputer Performance: Petrini, Kerbyson, Pakin, SC 2003. Operating-system noise measured at scale.
- Scientific Benchmarking of Parallel Computing Systems: Hoefler, Belli, SC 2015. Reporting rules for performance results.
- cgroup v2 admin guide and IRQ affinity: the kernel interfaces goset writes.
- fior512/goset: Go, Linux, Apache-2.0.