Step-by-step how-to: run two identical Redis instances on one Kubernetes node, one tuned for low latency by the balloons policy, the other left on normal CPUs, then compare the node configuration and benchmark both.
The balloons policy is configured in three groups of options:
- Control what a CPU is allowed to run
- no other container:
preferNewBalloons - no other kernel task: kernel
isolcpus+preferIsolCpus - no interrupts:
irqMode
- no other container:
- Configure CPUs
- hardware turbo priority:
pctPriority(Intel SST / Priority Core Turbo) - frequency locked at turbo:
minFreq=maxFreq=turbo,freqGovernor - wake-up latency:
disabledCstates
- hardware turbo priority:
- Configure container processes
- Linux scheduling policy and priority:
schedulingClass.policy,.priority - I/O scheduling class and priority:
schedulingClass.ioClass,.ioPriority
- Linux scheduling policy and priority:
Both CPU classes get the same frequency ceiling (turbo). The low-latency
class differs in its floor and governor, not in how fast it is allowed to go,
so the comparison is not decided by an artificial cap on the normal workload.
Results are intentionally not included. Measure in your own environment.
- Kubernetes 1.27+, one worker node,
kubectlandhelmworking. - Container runtime with NRI enabled: containerd 2.x (default on), containerd 1.7+ or CRI-O 1.26+ with NRI enabled.
- Node with ≥ 24 logical CPUs, of which ≥ 15 on the NUMA node holding the
isolated CPUs (2 + 2 + 8 for the three measured balloons, 1 for CPU 0 which is
left out, and enough left over for the background load of section 6.3). Scale
minCPUs/maxCPUsin section 3 down for smaller nodes. - Intel Xeon with Speed Select Technology (SST-CP + SST-TF), e.g. Xeon 6700P/6900P, for the Priority Core Turbo part. Without SST the PCT fields are ignored with a warning and everything else still works.
rootshell access on the node for verification (sudo).
isolcpus keeps the kernel scheduler from placing any task on the listed CPUs
unless the task is explicitly pinned there. Rules for picking them:
- at least 2 CPUs — the
minCPUsof the low-latency balloon type in section 3, - never CPU 0,
- all hyperthread siblings of every core you pick,
- 2 CPUs per socket, so the policy can place the balloon on either socket.
lscpu -e=CPU,CORE,SOCKET,NODE | head -20Set the CPUs to isolate, then add the boot parameters and reboot. The example below is for a 2-socket, 128-CPU node:
ISOLCPUS=1,2,65,66 # adjust to your node
sudo cp /etc/default/grub /etc/default/grub.bak
sudo sed -i "s/^\(GRUB_CMDLINE_LINUX_DEFAULT=\"[^\"]*\)\"/\1 isolcpus=managed_irq,domain,$ISOLCPUS nohz_full=$ISOLCPUS rcu_nocbs=$ISOLCPUS\"/" /etc/default/grub
grep GRUB_CMDLINE_LINUX_DEFAULT /etc/default/grub
sudo update-grub # RHEL/SUSE: sudo grub2-mkconfig -o /boot/grub2/grub.cfg
sudo rebootThat is the whole static configuration this how-to needs. Every parameter does something the balloons policy cannot do from user space:
isolcpus=domainremoves the CPUs from scheduler load balancing, so the kernel never places a runnable unpinned task there.isolcpus=managed_irqkeeps kernel-managed device queue interrupts (NVMe and NIC per-CPU queues) off them. Those are exactly the interrupts whosesmp_affinity_listis read-only, soirqModecannot move them.nohz_fullstops the periodic scheduler tick while a single task runs on a CPU and — more useful here — also defines the kernel's housekeeping mask for unbound kernel threads, keeping them off the isolated CPUs. Section 5.3 counts the difference this makes; section 6.4 measures what the tick suppression itself is worth for Redis.rcu_nocbsmoves RCU callback processing off the isolated CPUs.
One parameter commonly seen in realtime guides is deliberately not here:
irqaffinity=<housekeeping cpus>is redundant. It only sets the default affinity mask that IRQs get at boot. The balloons policy rewritessmp_affinity_listof every controllable IRQ at runtime throughirqMode, per balloon rather than once for the whole node, so the static mask adds nothing. Section 5.4 measures the result with and without it.
Verifying that the kernel command line has to stay this short is part of the exercise: section 5 measures every property it is supposed to deliver.
After reboot, give the node a minute to settle before measuring anything: right after boot there are still transient kworkers and the policy has not finished rewriting IRQ affinities, which inflates the counts in sections 5.3 and 5.4.
uptime -p
cat /proc/cmdline
cat /sys/devices/system/cpu/isolated
cat /sys/devices/system/cpu/nohz_fullchrt, ionice and lscpu come from util-linux. turbostat shows real
per-CPU frequency and C-state residency:
sudo apt-get install -y util-linux linux-tools-common linux-tools-$(uname -r)
sudo turbostat --versionintel-speed-select reads and writes SST state. It is not packaged by
distributions; build it from the kernel tree (33 MB sparse clone). It is only
needed for verifying and resetting SST — the balloons plugin talks to
/dev/isst_interface directly.
sudo apt-get install -y build-essential pkg-config libnl-3-dev libnl-genl-3-dev git
git clone --depth 1 --filter=blob:none --sparse https://github.com/torvalds/linux.git
cd linux
git sparse-checkout set tools/power/x86/intel-speed-select tools/include \
tools/build tools/scripts tools/lib include/uapi
make -C tools/power/x86/intel-speed-select
sudo make -C tools/power/x86/intel-speed-select install
cd ..
sudo intel-speed-select --infoExpected: SST-PP, SST-TF, SST-BF and SST-CP reported as supported.
--set allowPCT=true runs the plugin privileged and mounts host /dev at
/host/dev, which it needs to reach /dev/isst_interface for SST. Everything
else the policy writes (cpufreq and cpuidle sysfs, /proc/irq, cgroups) is
already available through the plugin's default host mounts. See
Installing with Helm.
helm repo add nri-plugins https://containers.github.io/nri-plugins
helm repo update nri-plugins
helm install nri-resource-policy-balloons \
nri-plugins/nri-resource-policy-balloons \
--namespace kube-system \
--set allowPCT=true
kubectl -n kube-system rollout status ds/nri-resource-policy-balloonsConfirm version v0.14.0 and the privileged security context:
kubectl -n kube-system get ds nri-resource-policy-balloons \
-o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}{.spec.template.spec.containers[0].securityContext}{"\n"}'Helm installed a BalloonsPolicy custom resource named default. Replace it
with a configuration that splits the node into housekeeping, normal and
low-latency CPUs. See
Managing Configuration with kubectl.
List the C-state names of your CPUs first — disabledCstates refers to them by
name:
grep . /sys/devices/system/cpu/cpu0/cpuidle/state*/name
grep . /sys/devices/system/cpu/cpu0/cpuidle/state*/latencyPick the NUMA node that holds the isolated CPUs. All three benchmarked balloons are pinned near it, so both Redis instances and the benchmark driver end up on the same socket and the measurement is not distorted by a cross-socket network path:
ISOL_FIRST=$(cut -d, -f1 /sys/devices/system/cpu/isolated | cut -d- -f1)
DEMO_NODE=$(basename /sys/devices/system/cpu/cpu$ISOL_FIRST/node*)
NODE_NCPU=$(ls -d /sys/devices/system/node/$DEMO_NODE/cpu[0-9]* | wc -l)
NOISE_CPUS=$(( NODE_NCPU - 14 )) # rest of the node, minus CPU 0, 2+2+8 and 1 spare
ORIG_GOV=$(cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor)
echo "isolated: $(cat /sys/devices/system/cpu/isolated) node: $DEMO_NODE ($NODE_NCPU CPUs)"
echo "noise balloon size: $NOISE_CPUS original cpufreq governor: $ORIG_GOV"NOISE_CPUS sizes the containerized background load of section 6.3, and
ORIG_GOV is restored by the reset policy in section 7. If NOISE_CPUS comes out
below 1, the NUMA node is too small: shrink the benchmark balloon or skip
section 6.3.
cat > low-latency-balloons.yaml <<EOF
apiVersion: config.nri/v1alpha1
kind: BalloonsPolicy
metadata:
name: default
namespace: kube-system
spec:
pinCPU: true
pinMemory: true
allocatorTopologyBalancing: false
availableResources:
# Leave CPU 0 out of every balloon: it carries unmovable timer and
# boot-time work, which would make it an unfair "normal" CPU here.
cpu: exclude-cpuset:0
reservedResources:
cpu: "2"
reservedPoolNamespaces:
- kube-system
balloonTypes:
# Low-latency workloads. One balloon per pod, on kernel-isolated CPUs,
# with interrupts moved away, HP turbo CPUs and realtime scheduling.
- name: low-latency
matchExpressions:
- key: pod/labels/workload-class
operator: Equals
values: ["low-latency"]
minCPUs: 2
maxCPUs: 2
preferNewBalloons: true # (1) no other container on these CPUs
preferIsolCpus: true # (1) no other kernel task on these CPUs
irqMode: isolate # (1) no unrelated IRQ on these CPUs
cpuClass: low-latency # (2) turbo + shallow C-states only
schedulingClass: low-latency # (3) SCHED_FIFO + realtime I/O priority
preferCloseToDevices: ["/sys/devices/system/node/$DEMO_NODE"]
# Normal workloads. Own CPUs, but default CPU and process tuning.
- name: normal
matchExpressions:
- key: pod/labels/workload-class
operator: Equals
values: ["normal"]
minCPUs: 2
maxCPUs: 2
preferNewBalloons: true
schedulingClass: normal
preferCloseToDevices: ["/sys/devices/system/node/$DEMO_NODE"]
# Benchmark driver. Sized so that the client never limits the result.
- name: benchmark
matchExpressions:
- key: pod/labels/workload-class
operator: Equals
values: ["benchmark"]
minCPUs: 8
maxCPUs: 8
preferNewBalloons: true
preferCloseToDevices: ["/sys/devices/system/node/$DEMO_NODE"]
# Containerized background load for section 6.3. Same CPU class and
# scheduling class as "normal", only bigger, and in its own balloon so
# that it never shares CPUs with redis-normal.
- name: noise
matchExpressions:
- key: pod/labels/workload-class
operator: Equals
values: ["noise"]
minCPUs: $NOISE_CPUS
maxCPUs: $NOISE_CPUS
preferNewBalloons: true
schedulingClass: normal
preferCloseToDevices: ["/sys/devices/system/node/$DEMO_NODE"]
# Housekeeping. Built-in "reserved" balloon takes kube-system pods,
# built-in "default" balloon takes every other pod. Both are IRQ sinks:
# all interrupts that no balloon claims land on their CPUs.
- name: reserved
irqMode: sink
- name: default
minBalloons: 1
minCPUs: 8
shareIdleCPUsInSame: system
irqMode: sink
cpuClasses:
# "default" is applied to idle CPUs and to every balloon type without an
# explicit cpuClass: full frequency range up to turbo, and "powersave",
# the governor a stock distribution leaves in place. Same ceiling as the
# low-latency class -- the normal workload is not capped.
- name: default
minFreq: min
maxFreq: turbo
freqGovernor: powersave
pctPriority: low
pctMinFreq: min
pctMaxFreq: turbo
# Locked at turbo: floor == ceiling, so there is no ramp-up after an idle
# period and no ramp-down between requests.
- name: low-latency
minFreq: turbo
maxFreq: turbo
freqGovernor: performance
disabledCstates: [C1E, C6, C6P] # adjust to your CPU's C-state names
pctPriority: high
pctMinFreq: turbo
pctMaxFreq: turbo
schedulingClasses:
- name: low-latency
policy: fifo
priority: 80
ioClass: rt
ioPriority: 0
- name: normal
policy: other
nice: 0
ioClass: be
ioPriority: 7
EOF
kubectl replace -f low-latency-balloons.yaml
kubectl -n kube-system rollout status ds/nri-resource-policy-balloonsGroup 1 — what a CPU is allowed to run
| Option | Effect | Reference |
|---|---|---|
matchExpressions |
routes pods to a balloon type by label | Choosing Balloon Type |
minCPUs, maxCPUs |
fixed balloon size, no inflation | Balloon Size Control |
preferNewBalloons: true |
every pod gets a fresh balloon, so no other container shares its CPUs | Choosing Balloon Instance |
preferIsolCpus: true |
allocate the balloon from kernel isolcpus, so the kernel scheduler places no other task there |
Static CPU Preferences |
irqMode: isolate |
remove these CPUs from the affinity of all IRQs that are neither claimed nor sinked | IRQ CPU Affinity Tuning |
irqMode: sink |
collect every unclaimed IRQ on the housekeeping CPUs | same |
shareIdleCPUsInSame: system |
housekeeping pods may burst onto any CPU that is in no balloon | Sharing idle CPUs |
noise balloon type |
containerized background load gets its own dedicated CPUs, so it competes for package power and cache but never for redis-normal's CPU time | Choosing Balloon Type |
reservedResources, reservedPoolNamespaces |
size and populate the built-in reserved balloon |
Reserved Balloon |
availableResources: cpu: exclude-cpuset:0 |
keep CPU 0 out of every balloon | Managed CPUs |
preferCloseToDevices |
place the balloon near a device — here the NUMA node of the isolated CPUs, so all benchmarked balloons share one socket | Static CPU Preferences |
Group 2 — CPU configuration (CPU Tuning)
| Option | Effect |
|---|---|
pctPriority: high |
Intel Priority Core Turbo: associate the CPUs to the SST-CP high-priority CLOS, and enable SST-TF so those cores keep the top turbo bucket when the package runs out of turbo budget. At most one high and one low class. See Priority Core Turbo (PCT). |
pctPriority: low |
low-priority CLOS. Required: idle and non-PCT CPUs fall back to it, so they do not consume the per-power-domain HP core budget. |
pctMinFreq, pctMaxFreq |
frequency window programmed into the CLOS. Both classes get pctMaxFreq: turbo here — the CLOS priority, not a cap, is what separates them under load. |
minFreq, maxFreq |
scaling_min_freq / scaling_max_freq in cpufreq sysfs. Symbolic values min, base, turbo are resolved from sysfs at runtime. minFreq: turbo on the low-latency class pins the floor to the ceiling. |
freqGovernor |
scaling_governor. With intel_pstate + HWP, performance makes the driver request max frequency and EPP performance continuously; powersave lets it scale between minFreq and maxFreq according to load and EPP. Section 6.4 measures what this is worth once the floor is already pinned. |
disabledCstates |
write 1 to cpuidle/state*/disable. Deep C-states have long exit latency; disabling them removes that latency from every wake-up. |
One hardware caveat on the "same ceiling" claim: while SST-TF is enabled (which
managed PCT mode does), the hardware itself limits LP cores to about the base
frequency, no matter what pctMaxFreq or maxFreq allow. That is what makes HP
cores special, so it cannot be configured away. Section 5.5 shows the effect:
the LP CPUs sit at base frequency under load even though their cpufreq limit reads
turbo. If you want a strictly equal frequency ceiling, drop pctPriority,
pctMinFreq and pctMaxFreq from both classes — the demo then compares only the
non-PCT mechanisms (frequency floor, governor, C-states, isolcpus, IRQs,
scheduling policy).
Group 3 — container process configuration (Scheduling and Priority)
| Option | Effect |
|---|---|
policy: fifo, priority: 80 |
sched_setscheduler(SCHED_FIFO, 80) on the container's processes: they are never preempted by SCHED_OTHER tasks |
ioClass: rt, ioPriority: 0 |
ioprio_set(IOPRIO_CLASS_RT, 0): highest block I/O priority |
policy: other, ioClass: be, ioPriority: 7 |
the normal workload, for comparison |
The scheduling class is applied when a container is created. Changing the policy configuration does not re-apply it to already running containers.
Two kernel constraints on realtime scheduling policies:
grep RT_GROUP_SCHED /boot/config-$(uname -r)
cat /proc/sys/kernel/sched_rt_runtime_us /proc/sys/kernel/sched_rt_period_us- With
CONFIG_RT_GROUP_SCHED=y, moving a task in a non-root cgroup to a realtime policy fails unless the cgroup has an RT runtime budget. Redis is fine on kernels built without it (Ubuntu, Debian, Fedora). sched_rt_runtime_us/sched_rt_period_uscap realtime tasks per CPU (950000/1000000 = 95 % by default,-1= uncapped). ASCHED_FIFOcontainer that busy-loops gets throttled at that limit. Redis blocks inepoll_wait, so it stays well below it.
Two identical Redis servers and one memtier_benchmark driver. Persistence is
off so that the measurement contains no disk I/O. CPU requests without
limits avoid CFS bandwidth throttling; balloons, not the CFS quota, decide
which CPUs the containers may use.
cat > workloads.yaml <<'EOF'
apiVersion: v1
kind: Namespace
metadata:
name: demo
---
apiVersion: v1
kind: Pod
metadata:
name: redis-lowlat
namespace: demo
labels:
app: redis-lowlat
workload-class: low-latency
spec:
containers:
- name: redis
image: docker.io/library/redis:7-alpine
args: ["--save", "", "--appendonly", "no", "--protected-mode", "no"]
resources:
requests:
cpu: "2"
memory: 1Gi
---
apiVersion: v1
kind: Service
metadata:
name: redis-lowlat
namespace: demo
spec:
selector:
app: redis-lowlat
ports:
- port: 6379
---
apiVersion: v1
kind: Pod
metadata:
name: redis-normal
namespace: demo
labels:
app: redis-normal
workload-class: normal
spec:
containers:
- name: redis
image: docker.io/library/redis:7-alpine
args: ["--save", "", "--appendonly", "no", "--protected-mode", "no"]
resources:
requests:
cpu: "2"
memory: 1Gi
---
apiVersion: v1
kind: Service
metadata:
name: redis-normal
namespace: demo
spec:
selector:
app: redis-normal
ports:
- port: 6379
---
apiVersion: v1
kind: Pod
metadata:
name: bench
namespace: demo
labels:
workload-class: benchmark
spec:
containers:
- name: memtier
image: docker.io/redislabs/memtier_benchmark:latest
command: ["sleep", "infinity"]
resources:
requests:
cpu: "8"
memory: 2Gi
---
apiVersion: v1
kind: Pod
metadata:
name: housekeeper
namespace: demo
spec:
containers:
- name: sleeper
image: docker.io/library/busybox:latest
command: ["sleep", "infinity"]
resources:
requests:
cpu: "500m"
memory: 64Mi
EOF
kubectl apply -f workloads.yaml
kubectl -n demo wait --for=condition=Ready pod --all --timeout=300shousekeeper carries no workload-class label, so it falls into the built-in
default balloon. See
Default Balloon.
Write the background-load pod too, but do not apply it yet — section 6.3 uses it.
One busy loop per CPU of its own balloon; nproc inside the container reports the
balloon's CPU count, so the container sizes itself:
cat > noise.yaml <<EOF
apiVersion: v1
kind: Pod
metadata:
name: noise
namespace: demo
labels:
workload-class: noise
spec:
containers:
- name: hog
image: docker.io/library/busybox:latest
command:
- sh
- -c
- 'echo "hogging \$(nproc) CPUs"; i=0; while [ \$i -lt \$(nproc) ]; do (while :; do :; done) & i=\$((i+1)); done; wait'
resources:
requests:
cpu: "$NOISE_CPUS"
memory: 64Mi
EOFThe policy exposes balloons as NodeResourceTopology zones
(NodeResourceTopology Integration):
kubectl get noderesourcetopologies -o json | jq -r '
["BALLOON","OWN CPUS","SHARED IDLE CPUS"],
(.items[].zones[] | select(.type=="balloon") | [
.name,
((.attributes[]|select(.name=="cpuset").value) // "-"),
((.attributes[]|select(.name=="shared cpuset").value) // "-")]) | @tsv' \
| column -t -s $'\t'Expect one low-latency[N] balloon, one normal[N], one benchmark[N], plus
reserved[0] for kube-system and default[0] for housekeeper. A built-in
balloon type that no container uses shows an empty cpuset.
low-latency, normal and benchmark must all sit inside $DEMO_NODE, and
CPU 0 must not appear in any of them:
cat /sys/devices/system/node/$DEMO_NODE/cpulistWhich balloon each container went into:
kubectl -n kube-system logs ds/nri-resource-policy-balloons \
| grep 'assigning container demo/' | tail -5Store the CPU sets for the verification commands below:
NRT='kubectl get noderesourcetopologies -o json'
bln_cpus() { eval $NRT | jq -r ".items[].zones[]
| select(.type==\"balloon\") | select(.name|startswith(\"$1\"))
| .attributes[] | select(.name==\"cpuset\").value" | head -1; }
LL_CPUS=$(bln_cpus low-latency)
NORM_CPUS=$(bln_cpus normal)
HK_CPUS=$(bln_cpus default)
echo "low-latency: $LL_CPUS normal: $NORM_CPUS housekeeping: $HK_CPUS"for p in redis-lowlat redis-normal bench housekeeper; do
printf '%-14s ' "$p"
kubectl -n demo exec $p -- grep -h -E 'Cpus_allowed_list|Mems_allowed_list' /proc/1/status | tr '\n\t' ' '
echo
doneThe two Redis containers must show disjoint CPU sets of equal size.
echo "kernel isolcpus: $(cat /sys/devices/system/cpu/isolated)"
echo "low-latency: $LL_CPUS"
echo "normal: $NORM_CPUS"$LL_CPUS must be a subset of isolated; $NORM_CPUS must not intersect it.
Count how many threads in the whole system are allowed to run on each CPU
(Cpus_allowed_list in /proc/<pid>/task/<tid>/status):
cat > cpulist.sh <<'EOF'
# expand_cpus "4-7,10" -> "4 5 6 7 10"
expand_cpus() {
local p
for p in ${1//,/ }; do
case $p in *-*) seq "${p%-*}" "${p#*-}" ;; *) echo "$p" ;; esac
done | tr '\n' ' '
}
EOF
cat > cpu-tasks.sh <<'EOF'
#!/bin/bash
# Usage: sudo ./cpu-tasks.sh CPULIST [-v]
. "$(dirname "$0")/cpulist.sh"
grep -H Cpus_allowed_list /proc/[0-9]*/task/[0-9]*/status 2>/dev/null |
awk -v want="$(expand_cpus "$1")" -v verbose="$2" '
function inlist(c,l, n,a,j,r){ n=split(l,a,","); for(j=1;j<=n;j++){
if(split(a[j],r,"-")==2){ if(c>=r[1]+0 && c<=r[2]+0) return 1 } else if(c+0==a[j]+0) return 1 } return 0 }
BEGIN{ m=split(want,w," ") }
{ split($1,p,"/"); pid=p[3]; tid=p[5]; mask=$2
cmd=""; f="/proc/" pid "/task/" tid "/comm"; if((getline l < f)>0) cmd=l; close(f)
for(i=1;i<=m;i++) if(inlist(w[i]+0,mask)) { cnt[w[i]]++; seen[w[i] SUBSEP cmd]=1 } }
END{ printf "%-5s %s\n", "CPU", "THREADS-ALLOWED"
for(i=1;i<=m;i++){ printf "%-5s %d\n", w[i], cnt[w[i]]+0
if(verbose=="-v") for(k in seen){ split(k,q,SUBSEP); if(q[1]==w[i]) printf " %s\n", q[2] } } }'
EOF
chmod +x cpu-tasks.sh
sudo ./cpu-tasks.sh "$LL_CPUS"
sudo ./cpu-tasks.sh "$NORM_CPUS"
sudo ./cpu-tasks.sh "$HK_CPUS"Expected: an order of magnitude fewer threads allowed on the low-latency CPUs.
Add -v to list them by name:
sudo ./cpu-tasks.sh "$LL_CPUS" -vWhat remains there are per-CPU kernel threads (kworker/<cpu>:*,
ksoftirqd/<cpu>, migration/<cpu>, cpuhp/<cpu>), unbound kworker rescuers
that only run when a workqueue stalls, and the low-latency pod itself.
This count is the clearest single measure of whether the kernel command line is
doing its job. isolcpus=domain alone does not shrink it much: it stops the load
balancer from migrating runnable tasks onto the CPU, but unbound kernel threads
still list it in their affinity mask and can be woken there. nohz_full is what
removes them, because the kernel derives its housekeeping mask for unbound
kthreads from it. Drop nohz_full from the command line and re-run this command
to see the difference on your node.
cat > irq-load.sh <<'EOF'
#!/bin/bash
# Usage: ./irq-load.sh CPULIST [SECONDS]
. "$(dirname "$0")/cpulist.sh"
cpus=$(expand_cpus "$1"); secs=${2:-5}
affinity_counts() { grep -h . /proc/irq/*/smp_affinity_list 2>/dev/null | awk '
{ n=split($0,p,",")
for(i=1;i<=n;i++) if(split(p[i],r,"-")==2){for(c=r[1];c<=r[2];c++)cnt[c]++} else cnt[p[i]]++ }
END{ for(c in cnt) print c, cnt[c] }'; }
delivered() { awk -v c=$1 'NR==1{for(i=1;i<=NF;i++) if($i=="CPU"c) col=i+1; next}
col && $col ~ /^[0-9]+$/ {s+=$col} END{print s+0}' /proc/interrupts; }
declare -A aff; while read -r c n; do aff[$c]=$n; done < <(affinity_counts)
declare -A t0; for c in $cpus; do t0[$c]=$(delivered $c); done
sleep "$secs"
printf '%-5s %-14s %s\n' CPU IRQS-AFFINE INTERRUPTS/s
for c in $cpus; do
printf '%-5s %-14s %s\n' "$c" "${aff[$c]:-0}" \
"$(( ( $(delivered $c) - ${t0[$c]} ) / secs ))"
done
EOF
chmod +x irq-load.sh
./irq-load.sh "$LL_CPUS,$NORM_CPUS,$HK_CPUS" 5Expected: the housekeeping CPUs are in the affinity mask of nearly every IRQ and receive nearly all interrupts; the low-latency CPUs are in very few and receive almost none.
The IRQs that remain on the low-latency CPUs are kernel-managed per-CPU device
queues (nvme*q*, seen as affinity is read-only warnings in the plugin log).
Their affinity cannot be changed from user space. They only fire for I/O
submitted from that CPU, which is what isolcpus=managed_irq limits.
Restrict which IRQs the policy touches with
controllableInterrupts,
for example controllableInterrupts: ["*eth0*", "*eno1*"], to silence those
warnings.
cat > cpu-config.sh <<'EOF'
#!/bin/bash
# Usage: sudo ./cpu-config.sh CPULIST
. "$(dirname "$0")/cpulist.sh"
printf '%-5s %-9s %-9s %-12s %-18s %s\n' CPU MIN-MHz MAX-MHz GOVERNOR DISABLED-CSTATES SST-CLOS
for cpu in $(expand_cpus "$1"); do
d=/sys/devices/system/cpu/cpu$cpu/cpufreq
dis=$(for s in /sys/devices/system/cpu/cpu$cpu/cpuidle/state*; do
[ "$(cat $s/disable)" = 1 ] && printf '%s ' "$(cat $s/name)"
done)
clos=$(intel-speed-select -c $cpu core-power get-assoc 2>&1 | sed -n 's/.*clos://p')
printf '%-5s %-9s %-9s %-12s %-18s %s\n' \
"$cpu" $(( $(cat $d/scaling_min_freq)/1000 )) $(( $(cat $d/scaling_max_freq)/1000 )) \
"$(cat $d/scaling_governor)" "${dis:--}" "${clos:--}"
done
EOF
chmod +x cpu-config.sh
sudo ./cpu-config.sh "$LL_CPUS,$NORM_CPUS,$HK_CPUS"Expected: low-latency CPUs at base..turbo MHz, in CLOS 0, with the deep
C-states disabled; all other CPUs at min..base MHz, in CLOS 3, all C-states
enabled.
SST-CP is enabled and CLOS-based (not proportional), and the two CLOSes carry
the frequency windows from pctMinFreq/pctMaxFreq:
sudo intel-speed-select core-power info 2>&1 | head -10
sudo intel-speed-select core-power get-config -c 0 2>&1 | head -11
sudo intel-speed-select core-power get-config -c 3 2>&1 | head -11The top turbo bucket that SST-TF grants to HP cores, and how many HP cores fit in it per power domain:
LEVEL=$(sudo intel-speed-select perf-profile get-config-current-level 2>&1 \
| sed -n 's/.*current_level://p' | head -1)
sudo intel-speed-select turbo-freq info -l "$LEVEL" 2>&1 | head -12Real frequency and C-state residency under load. Run both benchmarks at once and sample the node:
for s in redis-lowlat redis-normal; do
( kubectl -n demo exec bench -- memtier_benchmark -s $s -p 6379 \
--threads=2 --clients=4 --test-time=60 --data-size=32 \
--key-maximum=10000 --ratio=1:10 --hide-histogram > load-$s.log 2>&1 ) &
done
sleep 6
sudo turbostat --quiet --cpu "$LL_CPUS,$NORM_CPUS" \
--show CPU,Busy%,Avg_MHz,Bzy_MHz,C1%,C1E%,C6% --interval 8 --num_iterations 3
waitBzy_MHz is the frequency while the CPU was busy, C*% the residency in each
C-state. Read the rows where Busy% is near 100: each balloon has two CPUs and
Redis serves from one thread, so the other CPU of each balloon is idle. If no row
has a high Busy%, the sample missed the load — re-run the block.
Expect:
- low-latency CPU: an HP turbo frequency, matching one of the
high-priority-max-level-*-frequencyvalues from the previous command. Which level you get depends on how many cores are active node-wide, so it is not always the bucket-0 value, and it can differ between two runs of the same configuration. Treat sustained-throughput comparisons with that in mind. - normal CPU: about the base frequency — not the
turboits cpufreq limit reports. This is the SST-TF LP cap described above: while SST-TF is enabled, the hardware reserves the high turbo buckets for HP cores. - 0.00 residency in the C-states listed in
disabledCstates, on both the busy and the idle low-latency CPU. The normal balloon's idle CPU shows real C6 residency: that is the wake-up cost section 6.2 measures.
This step only needs to keep the server CPUs busy; ignore any "memtier may be the
bottleneck" warning in load-*.log here.
cat > proc-sched.sh <<'EOF'
#!/bin/bash
# Usage: sudo ./proc-sched.sh NAMESPACE POD
uid=$(kubectl -n "$1" get pod "$2" -o jsonpath='{.metadata.uid}')
printf '%-8s %-4s %-5s %-7s %-4s %-22s %s\n' TID CPU CLS RTPRIO NI IO-CLASS/PRIO COMMAND
for pid in $(grep -ls "pod$uid" /proc/[0-9]*/cgroup 2>/dev/null | cut -d/ -f3); do
while read -r tid cpu cls rt ni comm; do
[ "$comm" = pause ] && continue
printf '%-8s %-4s %-5s %-7s %-4s %-22s %s\n' "$tid" "$cpu" "$cls" "$rt" "$ni" \
"$(ionice -p "$tid" 2>/dev/null | tr -d '\n')" "$comm"
done < <(ps -Lo tid=,psr=,cls=,rtprio=,ni=,comm= -p "$pid" 2>/dev/null)
done
EOF
chmod +x proc-sched.sh
for p in redis-lowlat redis-normal; do echo "== $p"; sudo ./proc-sched.sh demo $p; doneExpected: every thread of redis-lowlat in class FF (SCHED_FIFO) at
RTPRIO 80 and I/O class realtime: prio 0; every thread of redis-normal in
class TS (SCHED_OTHER) at NI 0 and I/O class best-effort: prio 7.
CLS/RTPRIO/NI come from ps; the same values are readable per process
with chrt -p <tid> and ionice -p <tid>.
Two operating points. The first shows sustained throughput and tail latency, the second shows wake-up latency, where deep C-state exit and frequency ramp-up dominate.
memtier_benchmark reports Ops/sec and the Avg, p50, p99 and p99.9
latency of the whole request/response round trip, including the pod network path
that is identical for both servers.
for s in redis-lowlat redis-normal; do
echo "=== $s"
kubectl -n demo exec bench -- memtier_benchmark -s $s -p 6379 \
--threads=4 --clients=4 --test-time=30 --data-size=32 \
--key-maximum=10000 --ratio=1:10 --hide-histogram \
2>&1 | grep -E '^Type|^Totals|Cores used'
doneCores used shows how many CPUs of the benchmark balloon the driver consumed. If
it approaches --threads, raise --threads and, if needed, minCPUs/maxCPUs
of the benchmark balloon type, and repeat — otherwise the client, not the
server, is being measured. memtier_benchmark prints its own
"memtier may be the bottleneck" warning in that case.
--rate-limiting=200 sends 200 requests per second per client, so the server
CPU is idle between requests and every request pays the wake-up cost.
for s in redis-lowlat redis-normal; do
echo "=== $s"
kubectl -n demo exec bench -- memtier_benchmark -s $s -p 6379 \
--threads=1 --clients=1 --rate-limiting=200 --test-time=30 \
--data-size=32 --key-maximum=10000 --ratio=1:10 --hide-histogram \
2>&1 | grep -E '^Type|^Totals'
doneCompare Avg. Latency, p50, p99 and p99.9 between the two servers, and
compare C6 exit latency (/sys/devices/system/cpu/cpu0/cpuidle/state*/latency)
against the difference you measure.
In Kubernetes the neighbours are containers, so the background load is a pod. It
lands in the noise balloon type: same CPU class and scheduling class as
normal, its own dedicated CPUs, sized to fill the rest of $DEMO_NODE. Neither
Redis instance loses CPU time to it — both keep their dedicated CPUs. What the
noise does take away is package-level turbo budget, last-level cache and memory
bandwidth, which is exactly the contention pctPriority arbitrates.
kubectl apply -f noise.yaml
kubectl -n demo wait --for=condition=Ready pod/noise --timeout=120s
kubectl get noderesourcetopologies -o json | jq -r '
.items[].zones[] | select(.type=="balloon")
| [.name, ((.attributes[]|select(.name=="cpuset").value) // "-")] | @tsv' | column -t -s $'\t'The noise[0] cpuset must be disjoint from low-latency[0], normal[0] and
benchmark[0], and inside $DEMO_NODE.
Frequency under contention — this is where the HP and LP CLOSes separate:
NOISE_CPUS_ACTUAL=$(kubectl get noderesourcetopologies -o json | jq -r \
'.items[].zones[]|select(.name|startswith("noise"))|.attributes[]|select(.name=="cpuset").value')
sudo turbostat --quiet --cpu "$LL_CPUS,$NORM_CPUS,$NOISE_CPUS_ACTUAL" \
--show CPU,Busy%,Bzy_MHz --interval 8 --num_iterations 2Repeat both measurements:
for s in redis-lowlat redis-normal; do
echo "=== $s, sustained, with noise"
kubectl -n demo exec bench -- memtier_benchmark -s $s -p 6379 \
--threads=4 --clients=4 --test-time=30 --data-size=32 \
--key-maximum=10000 --ratio=1:10 --hide-histogram \
2>&1 | grep -E '^Type|^Totals|Cores used'
done
for s in redis-lowlat redis-normal; do
echo "=== $s, 200 req/s, with noise"
kubectl -n demo exec bench -- memtier_benchmark -s $s -p 6379 \
--threads=1 --clients=1 --rate-limiting=200 --test-time=30 \
--data-size=32 --key-maximum=10000 --ratio=1:10 --hide-histogram \
2>&1 | grep -E '^Type|^Totals'
doneRemove the load again:
kubectl delete -f noise.yamlTwo effects to separate when reading the numbers:
- The low-latency CPUs sit in the HP CLOS with
pctMinFreq: turbo, so their frequency should hold while the LP cores around them lose turbo headroom. - A busy neighbour also keeps its own CPU out of deep C-states. It cannot do that for the normal Redis CPU, because that CPU is dedicated and still idles between requests — so the low request rate gap does not close.
Three configuration choices from section 1.1 and section 3 that only measurement can settle. Run 6.1 and 6.2 again after each change and compare.
irqaffinity — add irqaffinity=<every CPU except $ISOLCPUS> to the kernel
command line, reboot, and re-run section 5.4 rather than the benchmarks: the
question is whether the static boot mask reaches any interrupt that irqMode does
not.
./irq-load.sh "$LL_CPUS,$NORM_CPUS,$HK_CPUS" 5Compare both columns. A lower IRQS-AFFINE on the low-latency CPUs with no change
in INTERRUPTS/s means the static mask only narrowed interrupts that never fire.
nohz_full — reboot with nohz_full=$ISOLCPUS removed from the kernel
command line, keeping everything else identical:
cat /sys/devices/system/cpu/nohz_full # the isolated CPUs, or empty
sudo ./cpu-tasks.sh "$LL_CPUS" # compare against section 5.3Two effects pull in opposite directions, so measure both columns:
- A
nohz_fullCPU skips the periodic tick while a single task runs, but pays extra tick and RCU bookkeeping on every user/kernel transition. Redis crosses that boundary several times per request, so tick suppression is not obviously a win for it. - Without
nohz_fullthe kernel has no housekeeping mask for unbound kernel threads, and the number of threads allowed on the low-latency CPUs grows by roughly a factor of three. Those threads rarely run, but nothing stops them.
freqGovernor — the low-latency class already pins minFreq to turbo. Test
whether the governor still matters on top of that by switching the low-latency
class between performance and powersave:
llclass() { # patch freqGovernor of the low-latency cpuClass in place
local idx
idx=$(kubectl -n kube-system get balloonspolicy default -o json \
| jq '[.spec.cpuClasses[].name] | index("low-latency")')
kubectl -n kube-system patch balloonspolicy default --type=json \
-p "[{\"op\":\"replace\",\"path\":\"/spec/cpuClasses/$idx/freqGovernor\",\"value\":\"$1\"}]"
}
llclass powersave
sleep 10
sudo ./cpu-config.sh "$LL_CPUS" # governor changed, min and max still turbo?
# repeat 6.1 and 6.2 here
llclass performance
sleep 10
sudo ./cpu-config.sh "$LL_CPUS"The jq index lookup keeps the patch on the low-latency class only; editing the
YAML with sed would also hit the default class, which must stay on
powersave.
With SST-TF enabled the hardware caps LP cores near base frequency, so part of the sustained-throughput gap is the CLOS priority rather than anything else in the configuration. Separate the two by taking PCT out entirely and repeating 6.1 and 6.2:
sed '/^ *pct\(Priority\|MinFreq\|MaxFreq\):/d' low-latency-balloons.yaml \
> no-pct-balloons.yaml
diff low-latency-balloons.yaml no-pct-balloons.yaml
kubectl replace -f no-pct-balloons.yaml
sleep 12
# the policy does not undo SoC-wide SST state when PCT disappears from the config
sudo intel-speed-select core-power disable
sudo intel-speed-select turbo-freq disable
sudo intel-speed-select core-power info 2>&1 | head -9 # enable-status:disabled
sudo ./cpu-config.sh "$LL_CPUS,$NORM_CPUS" # cpufreq limits unchanged
# repeat 6.1 and 6.2 hereNow both classes really do have the same frequency ceiling, and the remaining
low-latency advantage comes only from the pinned frequency floor, the disabled
C-states, isolcpus, the IRQ isolation and the realtime scheduling policy.
The SST-CLOS column of cpu-config.sh keeps reporting the last association even
after core-power disable; core-power info is the field that tells you whether
CLOSes are in effect at all.
Put PCT back afterwards:
kubectl replace -f low-latency-balloons.yaml
sleep 12
sudo intel-speed-select core-power info 2>&1 | head -9 # enable-status:enabled
sudo ./cpu-config.sh "$LL_CPUS,$NORM_CPUS" # CLOS 0 and 3 againOrder matters. Delete the workloads first: scheduling policy and I/O priority are set at container creation and are not reverted by a configuration change.
kubectl delete -f noise.yaml --ignore-not-found
kubectl delete -f workloads.yamlApply a reset policy that puts every container into one balloon covering all CPUs, restores cpufreq limits and C-states, spreads IRQs back over all CPUs, and returns SST to CLOS 0 with the full frequency range. This mirrors Reset CPU and Memory Pinning.
N=$(nproc --all)
cat > reset-balloons.yaml <<EOF
apiVersion: config.nri/v1alpha1
kind: BalloonsPolicy
metadata:
name: default
namespace: kube-system
spec:
pinCPU: true
pinMemory: true
reservedResources:
cpu: cpuset:0-$((N-1))
balloonTypes:
- name: reserved
namespaces: ["*"] # every container into this balloon
minCPUs: $N # balloon covers every CPU
maxCPUs: $N
preferIsolCpus: true # including the kernel-isolated ones
irqClaim: ["*"] # every IRQ back onto every CPU
cpuClasses:
# Applied to balloon CPUs and to idle CPUs: full frequency range, the
# governor the node had before the demo, no C-state disabled, all CPUs
# into SST CLOS 0 with no cap.
- name: default
minFreq: min
maxFreq: turbo
freqGovernor: $ORIG_GOV
pctPriority: high
pctMinFreq: min
pctMaxFreq: turbo
EOF
kubectl replace -f reset-balloons.yaml
sleep 20Verify the reset:
sudo ./cpu-config.sh 0,1,2,$((N-1)) # min..turbo, no disabled C-state, CLOS 0
./irq-load.sh 0,1,2,$((N-1)) 5 # every CPU in nearly every IRQ affinity
sudo intel-speed-select core-power get-config -c 0 2>&1 | head -11
kubectl get noderesourcetopologies -o json | jq -r \
'.items[].zones[]|select(.type=="balloon")|[.name,((.attributes[]|select(.name=="cpuset").value)//"-")]|@tsv'
# cpusets of every running container: only "0-<N-1>" (and empty) must be left
find /sys/fs/cgroup/kubepods* -name cpuset.cpus -print0 | xargs -0 cat | sort -uUninstall. The plugin does not turn SST off on exit, so disable SST-CP and
SST-TF explicitly — this restores enable-status:disabled,
clos-enable-status:disabled, priority-type:proportional. See
Uninstalling with Helm.
helm uninstall nri-resource-policy-balloons -n kube-system
kubectl delete crd balloonspolicies.config.nri
sudo intel-speed-select core-power disable
sudo intel-speed-select turbo-freq disable
sudo intel-speed-select core-power info 2>&1 | head -10Final check that nothing is left tuned:
grep -h . /sys/devices/system/cpu/cpu*/cpufreq/scaling_min_freq | sort -u
grep -h . /sys/devices/system/cpu/cpu*/cpufreq/scaling_max_freq | sort -u
grep -h . /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor | sort -u # $ORIG_GOV
grep -l 1 /sys/devices/system/cpu/cpu*/cpuidle/state*/disable | wc -l # 0If the last count is not 0, the plugin pod was restarted between disabling the
C-states and applying the reset policy: its cpuidle writer is created lazily, so
a fresh process with no disabledCstates anywhere never touches cpuidle sysfs.
Re-enable them directly:
for f in $(grep -l 1 /sys/devices/system/cpu/cpu*/cpuidle/state*/disable); do
echo 0 | sudo tee "$f" > /dev/null
done
grep -l 1 /sys/devices/system/cpu/cpu*/cpuidle/state*/disable | wc -l # 0Two things are not restored by these commands:
- The exact boot-time IRQ affinity from the
irqaffinity=kernel parameter. The reset policy spreads IRQs over all CPUs, including the isolated ones. A reboot restores the boot-time masks. - The
isolcpus,nohz_fullandrcu_nocbskernel parameters. Remove them from/etc/default/grub(or restore/etc/default/grub.bak), runsudo update-gruband reboot.
sudo cp /etc/default/grub.bak /etc/default/grub && sudo update-grub && sudo reboot- Balloons policy documentation
- Installing with Helm
- Managing Configuration with kubectl
- Container-to-Balloon Assignment
- CPUs-to-Balloon Selection
- Static CPU Preferences —
preferIsolCpus - Balloon Size Control —
minCPUs,maxCPUs,minBalloons - Scheduling and Priority —
schedulingClasses - Sharing idle CPUs
- CPU Tuning —
cpuClasses - Priority Core Turbo (PCT) —
pctPriority - IRQ CPU Affinity Tuning —
irqMode,irqClaim - Restricting IRQ CPU Affinity Tuning —
controllableInterrupts - Built-in Balloon Types —
reserved,default - Managed CPUs —
availableResources - Reset CPU and Memory Pinning
- Latency-Critical Containers — cookbook
- Troubleshooting
- Intel Speed Select Technology in the Linux kernel
sched_setscheduler(2),ioprio_set(2)- memtier_benchmark