Skip to content

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

Low-latency and realtime tuning with the NRI balloons policy

Step-by-step how-to: run two identical Redis instances on one Kubernetes node, one tuned for low latency by the balloons policy, the other left on normal CPUs, then compare the node configuration and benchmark both.

The balloons policy is configured in three groups of options:

  1. Control what a CPU is allowed to run
    • no other container: preferNewBalloons
    • no other kernel task: kernel isolcpus + preferIsolCpus
    • no interrupts: irqMode
  2. Configure CPUs
    • hardware turbo priority: pctPriority (Intel SST / Priority Core Turbo)
    • frequency locked at turbo: minFreq = maxFreq = turbo, freqGovernor
    • wake-up latency: disabledCstates
  3. Configure container processes
    • Linux scheduling policy and priority: schedulingClass.policy, .priority
    • I/O scheduling class and priority: schedulingClass.ioClass, .ioPriority

Both CPU classes get the same frequency ceiling (turbo). The low-latency class differs in its floor and governor, not in how fast it is allowed to go, so the comparison is not decided by an artificial cap on the normal workload.

Results are intentionally not included. Measure in your own environment.


1. Prerequisites

  • Kubernetes 1.27+, one worker node, kubectl and helm working.
  • Container runtime with NRI enabled: containerd 2.x (default on), containerd 1.7+ or CRI-O 1.26+ with NRI enabled.
  • Node with ≥ 24 logical CPUs, of which ≥ 15 on the NUMA node holding the isolated CPUs (2 + 2 + 8 for the three measured balloons, 1 for CPU 0 which is left out, and enough left over for the background load of section 6.3). Scale minCPUs/maxCPUs in section 3 down for smaller nodes.
  • Intel Xeon with Speed Select Technology (SST-CP + SST-TF), e.g. Xeon 6700P/6900P, for the Priority Core Turbo part. Without SST the PCT fields are ignored with a warning and everything else still works.
  • root shell access on the node for verification (sudo).

1.1 Isolate CPUs on the kernel command line

isolcpus keeps the kernel scheduler from placing any task on the listed CPUs unless the task is explicitly pinned there. Rules for picking them:

  • at least 2 CPUs — the minCPUs of the low-latency balloon type in section 3,
  • never CPU 0,
  • all hyperthread siblings of every core you pick,
  • 2 CPUs per socket, so the policy can place the balloon on either socket.
lscpu -e=CPU,CORE,SOCKET,NODE | head -20

Set the CPUs to isolate, then add the boot parameters and reboot. The example below is for a 2-socket, 128-CPU node:

ISOLCPUS=1,2,65,66     # adjust to your node

sudo cp /etc/default/grub /etc/default/grub.bak
sudo sed -i "s/^\(GRUB_CMDLINE_LINUX_DEFAULT=\"[^\"]*\)\"/\1 isolcpus=managed_irq,domain,$ISOLCPUS nohz_full=$ISOLCPUS rcu_nocbs=$ISOLCPUS\"/" /etc/default/grub
grep GRUB_CMDLINE_LINUX_DEFAULT /etc/default/grub
sudo update-grub    # RHEL/SUSE: sudo grub2-mkconfig -o /boot/grub2/grub.cfg
sudo reboot

That is the whole static configuration this how-to needs. Every parameter does something the balloons policy cannot do from user space:

  • isolcpus=domain removes the CPUs from scheduler load balancing, so the kernel never places a runnable unpinned task there.
  • isolcpus=managed_irq keeps kernel-managed device queue interrupts (NVMe and NIC per-CPU queues) off them. Those are exactly the interrupts whose smp_affinity_list is read-only, so irqMode cannot move them.
  • nohz_full stops the periodic scheduler tick while a single task runs on a CPU and — more useful here — also defines the kernel's housekeeping mask for unbound kernel threads, keeping them off the isolated CPUs. Section 5.3 counts the difference this makes; section 6.4 measures what the tick suppression itself is worth for Redis.
  • rcu_nocbs moves RCU callback processing off the isolated CPUs.

One parameter commonly seen in realtime guides is deliberately not here:

  • irqaffinity=<housekeeping cpus> is redundant. It only sets the default affinity mask that IRQs get at boot. The balloons policy rewrites smp_affinity_list of every controllable IRQ at runtime through irqMode, per balloon rather than once for the whole node, so the static mask adds nothing. Section 5.4 measures the result with and without it.

Verifying that the kernel command line has to stay this short is part of the exercise: section 5 measures every property it is supposed to deliver.

After reboot, give the node a minute to settle before measuring anything: right after boot there are still transient kworkers and the policy has not finished rewriting IRQ affinities, which inflates the counts in sections 5.3 and 5.4.

uptime -p
cat /proc/cmdline
cat /sys/devices/system/cpu/isolated
cat /sys/devices/system/cpu/nohz_full

1.2 Node tools used for verification

chrt, ionice and lscpu come from util-linux. turbostat shows real per-CPU frequency and C-state residency:

sudo apt-get install -y util-linux linux-tools-common linux-tools-$(uname -r)
sudo turbostat --version

intel-speed-select reads and writes SST state. It is not packaged by distributions; build it from the kernel tree (33 MB sparse clone). It is only needed for verifying and resetting SST — the balloons plugin talks to /dev/isst_interface directly.

sudo apt-get install -y build-essential pkg-config libnl-3-dev libnl-genl-3-dev git

git clone --depth 1 --filter=blob:none --sparse https://github.com/torvalds/linux.git
cd linux
git sparse-checkout set tools/power/x86/intel-speed-select tools/include \
    tools/build tools/scripts tools/lib include/uapi
make -C tools/power/x86/intel-speed-select
sudo make -C tools/power/x86/intel-speed-select install
cd ..

sudo intel-speed-select --info

Expected: SST-PP, SST-TF, SST-BF and SST-CP reported as supported.


2. Install the balloons policy

--set allowPCT=true runs the plugin privileged and mounts host /dev at /host/dev, which it needs to reach /dev/isst_interface for SST. Everything else the policy writes (cpufreq and cpuidle sysfs, /proc/irq, cgroups) is already available through the plugin's default host mounts. See Installing with Helm.

helm repo add nri-plugins https://containers.github.io/nri-plugins
helm repo update nri-plugins

helm install nri-resource-policy-balloons \
    nri-plugins/nri-resource-policy-balloons \
    --namespace kube-system \
    --set allowPCT=true

kubectl -n kube-system rollout status ds/nri-resource-policy-balloons

Confirm version v0.14.0 and the privileged security context:

kubectl -n kube-system get ds nri-resource-policy-balloons \
    -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}{.spec.template.spec.containers[0].securityContext}{"\n"}'

3. Replace the default policy configuration

Helm installed a BalloonsPolicy custom resource named default. Replace it with a configuration that splits the node into housekeeping, normal and low-latency CPUs. See Managing Configuration with kubectl.

List the C-state names of your CPUs first — disabledCstates refers to them by name:

grep . /sys/devices/system/cpu/cpu0/cpuidle/state*/name
grep . /sys/devices/system/cpu/cpu0/cpuidle/state*/latency

Pick the NUMA node that holds the isolated CPUs. All three benchmarked balloons are pinned near it, so both Redis instances and the benchmark driver end up on the same socket and the measurement is not distorted by a cross-socket network path:

ISOL_FIRST=$(cut -d, -f1 /sys/devices/system/cpu/isolated | cut -d- -f1)
DEMO_NODE=$(basename /sys/devices/system/cpu/cpu$ISOL_FIRST/node*)
NODE_NCPU=$(ls -d /sys/devices/system/node/$DEMO_NODE/cpu[0-9]* | wc -l)
NOISE_CPUS=$(( NODE_NCPU - 14 ))       # rest of the node, minus CPU 0, 2+2+8 and 1 spare
ORIG_GOV=$(cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor)
echo "isolated: $(cat /sys/devices/system/cpu/isolated)  node: $DEMO_NODE ($NODE_NCPU CPUs)"
echo "noise balloon size: $NOISE_CPUS  original cpufreq governor: $ORIG_GOV"

NOISE_CPUS sizes the containerized background load of section 6.3, and ORIG_GOV is restored by the reset policy in section 7. If NOISE_CPUS comes out below 1, the NUMA node is too small: shrink the benchmark balloon or skip section 6.3.

cat > low-latency-balloons.yaml <<EOF
apiVersion: config.nri/v1alpha1
kind: BalloonsPolicy
metadata:
  name: default
  namespace: kube-system
spec:
  pinCPU: true
  pinMemory: true
  allocatorTopologyBalancing: false
  availableResources:
    # Leave CPU 0 out of every balloon: it carries unmovable timer and
    # boot-time work, which would make it an unfair "normal" CPU here.
    cpu: exclude-cpuset:0
  reservedResources:
    cpu: "2"
  reservedPoolNamespaces:
  - kube-system

  balloonTypes:

  # Low-latency workloads. One balloon per pod, on kernel-isolated CPUs,
  # with interrupts moved away, HP turbo CPUs and realtime scheduling.
  - name: low-latency
    matchExpressions:
    - key: pod/labels/workload-class
      operator: Equals
      values: ["low-latency"]
    minCPUs: 2
    maxCPUs: 2
    preferNewBalloons: true        # (1) no other container on these CPUs
    preferIsolCpus: true           # (1) no other kernel task on these CPUs
    irqMode: isolate               # (1) no unrelated IRQ on these CPUs
    cpuClass: low-latency          # (2) turbo + shallow C-states only
    schedulingClass: low-latency   # (3) SCHED_FIFO + realtime I/O priority
    preferCloseToDevices: ["/sys/devices/system/node/$DEMO_NODE"]

  # Normal workloads. Own CPUs, but default CPU and process tuning.
  - name: normal
    matchExpressions:
    - key: pod/labels/workload-class
      operator: Equals
      values: ["normal"]
    minCPUs: 2
    maxCPUs: 2
    preferNewBalloons: true
    schedulingClass: normal
    preferCloseToDevices: ["/sys/devices/system/node/$DEMO_NODE"]

  # Benchmark driver. Sized so that the client never limits the result.
  - name: benchmark
    matchExpressions:
    - key: pod/labels/workload-class
      operator: Equals
      values: ["benchmark"]
    minCPUs: 8
    maxCPUs: 8
    preferNewBalloons: true
    preferCloseToDevices: ["/sys/devices/system/node/$DEMO_NODE"]

  # Containerized background load for section 6.3. Same CPU class and
  # scheduling class as "normal", only bigger, and in its own balloon so
  # that it never shares CPUs with redis-normal.
  - name: noise
    matchExpressions:
    - key: pod/labels/workload-class
      operator: Equals
      values: ["noise"]
    minCPUs: $NOISE_CPUS
    maxCPUs: $NOISE_CPUS
    preferNewBalloons: true
    schedulingClass: normal
    preferCloseToDevices: ["/sys/devices/system/node/$DEMO_NODE"]

  # Housekeeping. Built-in "reserved" balloon takes kube-system pods,
  # built-in "default" balloon takes every other pod. Both are IRQ sinks:
  # all interrupts that no balloon claims land on their CPUs.
  - name: reserved
    irqMode: sink
  - name: default
    minBalloons: 1
    minCPUs: 8
    shareIdleCPUsInSame: system
    irqMode: sink

  cpuClasses:
  # "default" is applied to idle CPUs and to every balloon type without an
  # explicit cpuClass: full frequency range up to turbo, and "powersave",
  # the governor a stock distribution leaves in place. Same ceiling as the
  # low-latency class -- the normal workload is not capped.
  - name: default
    minFreq: min
    maxFreq: turbo
    freqGovernor: powersave
    pctPriority: low
    pctMinFreq: min
    pctMaxFreq: turbo
  # Locked at turbo: floor == ceiling, so there is no ramp-up after an idle
  # period and no ramp-down between requests.
  - name: low-latency
    minFreq: turbo
    maxFreq: turbo
    freqGovernor: performance
    disabledCstates: [C1E, C6, C6P]   # adjust to your CPU's C-state names
    pctPriority: high
    pctMinFreq: turbo
    pctMaxFreq: turbo

  schedulingClasses:
  - name: low-latency
    policy: fifo
    priority: 80
    ioClass: rt
    ioPriority: 0
  - name: normal
    policy: other
    nice: 0
    ioClass: be
    ioPriority: 7
EOF

kubectl replace -f low-latency-balloons.yaml
kubectl -n kube-system rollout status ds/nri-resource-policy-balloons

What each option does

Group 1 — what a CPU is allowed to run

Option Effect Reference
matchExpressions routes pods to a balloon type by label Choosing Balloon Type
minCPUs, maxCPUs fixed balloon size, no inflation Balloon Size Control
preferNewBalloons: true every pod gets a fresh balloon, so no other container shares its CPUs Choosing Balloon Instance
preferIsolCpus: true allocate the balloon from kernel isolcpus, so the kernel scheduler places no other task there Static CPU Preferences
irqMode: isolate remove these CPUs from the affinity of all IRQs that are neither claimed nor sinked IRQ CPU Affinity Tuning
irqMode: sink collect every unclaimed IRQ on the housekeeping CPUs same
shareIdleCPUsInSame: system housekeeping pods may burst onto any CPU that is in no balloon Sharing idle CPUs
noise balloon type containerized background load gets its own dedicated CPUs, so it competes for package power and cache but never for redis-normal's CPU time Choosing Balloon Type
reservedResources, reservedPoolNamespaces size and populate the built-in reserved balloon Reserved Balloon
availableResources: cpu: exclude-cpuset:0 keep CPU 0 out of every balloon Managed CPUs
preferCloseToDevices place the balloon near a device — here the NUMA node of the isolated CPUs, so all benchmarked balloons share one socket Static CPU Preferences

Group 2 — CPU configuration (CPU Tuning)

Option Effect
pctPriority: high Intel Priority Core Turbo: associate the CPUs to the SST-CP high-priority CLOS, and enable SST-TF so those cores keep the top turbo bucket when the package runs out of turbo budget. At most one high and one low class. See Priority Core Turbo (PCT).
pctPriority: low low-priority CLOS. Required: idle and non-PCT CPUs fall back to it, so they do not consume the per-power-domain HP core budget.
pctMinFreq, pctMaxFreq frequency window programmed into the CLOS. Both classes get pctMaxFreq: turbo here — the CLOS priority, not a cap, is what separates them under load.
minFreq, maxFreq scaling_min_freq / scaling_max_freq in cpufreq sysfs. Symbolic values min, base, turbo are resolved from sysfs at runtime. minFreq: turbo on the low-latency class pins the floor to the ceiling.
freqGovernor scaling_governor. With intel_pstate + HWP, performance makes the driver request max frequency and EPP performance continuously; powersave lets it scale between minFreq and maxFreq according to load and EPP. Section 6.4 measures what this is worth once the floor is already pinned.
disabledCstates write 1 to cpuidle/state*/disable. Deep C-states have long exit latency; disabling them removes that latency from every wake-up.

One hardware caveat on the "same ceiling" claim: while SST-TF is enabled (which managed PCT mode does), the hardware itself limits LP cores to about the base frequency, no matter what pctMaxFreq or maxFreq allow. That is what makes HP cores special, so it cannot be configured away. Section 5.5 shows the effect: the LP CPUs sit at base frequency under load even though their cpufreq limit reads turbo. If you want a strictly equal frequency ceiling, drop pctPriority, pctMinFreq and pctMaxFreq from both classes — the demo then compares only the non-PCT mechanisms (frequency floor, governor, C-states, isolcpus, IRQs, scheduling policy).

Group 3 — container process configuration (Scheduling and Priority)

Option Effect
policy: fifo, priority: 80 sched_setscheduler(SCHED_FIFO, 80) on the container's processes: they are never preempted by SCHED_OTHER tasks
ioClass: rt, ioPriority: 0 ioprio_set(IOPRIO_CLASS_RT, 0): highest block I/O priority
policy: other, ioClass: be, ioPriority: 7 the normal workload, for comparison

The scheduling class is applied when a container is created. Changing the policy configuration does not re-apply it to already running containers.

Two kernel constraints on realtime scheduling policies:

grep RT_GROUP_SCHED /boot/config-$(uname -r)
cat /proc/sys/kernel/sched_rt_runtime_us /proc/sys/kernel/sched_rt_period_us
  • With CONFIG_RT_GROUP_SCHED=y, moving a task in a non-root cgroup to a realtime policy fails unless the cgroup has an RT runtime budget. Redis is fine on kernels built without it (Ubuntu, Debian, Fedora).
  • sched_rt_runtime_us/sched_rt_period_us cap realtime tasks per CPU (950000/1000000 = 95 % by default, -1 = uncapped). A SCHED_FIFO container that busy-loops gets throttled at that limit. Redis blocks in epoll_wait, so it stays well below it.

4. Deploy the workloads

Two identical Redis servers and one memtier_benchmark driver. Persistence is off so that the measurement contains no disk I/O. CPU requests without limits avoid CFS bandwidth throttling; balloons, not the CFS quota, decide which CPUs the containers may use.

cat > workloads.yaml <<'EOF'
apiVersion: v1
kind: Namespace
metadata:
  name: demo
---
apiVersion: v1
kind: Pod
metadata:
  name: redis-lowlat
  namespace: demo
  labels:
    app: redis-lowlat
    workload-class: low-latency
spec:
  containers:
  - name: redis
    image: docker.io/library/redis:7-alpine
    args: ["--save", "", "--appendonly", "no", "--protected-mode", "no"]
    resources:
      requests:
        cpu: "2"
        memory: 1Gi
---
apiVersion: v1
kind: Service
metadata:
  name: redis-lowlat
  namespace: demo
spec:
  selector:
    app: redis-lowlat
  ports:
  - port: 6379
---
apiVersion: v1
kind: Pod
metadata:
  name: redis-normal
  namespace: demo
  labels:
    app: redis-normal
    workload-class: normal
spec:
  containers:
  - name: redis
    image: docker.io/library/redis:7-alpine
    args: ["--save", "", "--appendonly", "no", "--protected-mode", "no"]
    resources:
      requests:
        cpu: "2"
        memory: 1Gi
---
apiVersion: v1
kind: Service
metadata:
  name: redis-normal
  namespace: demo
spec:
  selector:
    app: redis-normal
  ports:
  - port: 6379
---
apiVersion: v1
kind: Pod
metadata:
  name: bench
  namespace: demo
  labels:
    workload-class: benchmark
spec:
  containers:
  - name: memtier
    image: docker.io/redislabs/memtier_benchmark:latest
    command: ["sleep", "infinity"]
    resources:
      requests:
        cpu: "8"
        memory: 2Gi
---
apiVersion: v1
kind: Pod
metadata:
  name: housekeeper
  namespace: demo
spec:
  containers:
  - name: sleeper
    image: docker.io/library/busybox:latest
    command: ["sleep", "infinity"]
    resources:
      requests:
        cpu: "500m"
        memory: 64Mi
EOF

kubectl apply -f workloads.yaml
kubectl -n demo wait --for=condition=Ready pod --all --timeout=300s

housekeeper carries no workload-class label, so it falls into the built-in default balloon. See Default Balloon.

Write the background-load pod too, but do not apply it yet — section 6.3 uses it. One busy loop per CPU of its own balloon; nproc inside the container reports the balloon's CPU count, so the container sizes itself:

cat > noise.yaml <<EOF
apiVersion: v1
kind: Pod
metadata:
  name: noise
  namespace: demo
  labels:
    workload-class: noise
spec:
  containers:
  - name: hog
    image: docker.io/library/busybox:latest
    command:
    - sh
    - -c
    - 'echo "hogging \$(nproc) CPUs"; i=0; while [ \$i -lt \$(nproc) ]; do (while :; do :; done) & i=\$((i+1)); done; wait'
    resources:
      requests:
        cpu: "$NOISE_CPUS"
        memory: 64Mi
EOF

5. Verify the configuration

5.1 Balloons and their CPUs

The policy exposes balloons as NodeResourceTopology zones (NodeResourceTopology Integration):

kubectl get noderesourcetopologies -o json | jq -r '
  ["BALLOON","OWN CPUS","SHARED IDLE CPUS"],
  (.items[].zones[] | select(.type=="balloon") | [
     .name,
     ((.attributes[]|select(.name=="cpuset").value) // "-"),
     ((.attributes[]|select(.name=="shared cpuset").value) // "-")]) | @tsv' \
  | column -t -s $'\t'

Expect one low-latency[N] balloon, one normal[N], one benchmark[N], plus reserved[0] for kube-system and default[0] for housekeeper. A built-in balloon type that no container uses shows an empty cpuset.

low-latency, normal and benchmark must all sit inside $DEMO_NODE, and CPU 0 must not appear in any of them:

cat /sys/devices/system/node/$DEMO_NODE/cpulist

Which balloon each container went into:

kubectl -n kube-system logs ds/nri-resource-policy-balloons \
    | grep 'assigning container demo/' | tail -5

Store the CPU sets for the verification commands below:

NRT='kubectl get noderesourcetopologies -o json'
bln_cpus() { eval $NRT | jq -r ".items[].zones[]
    | select(.type==\"balloon\") | select(.name|startswith(\"$1\"))
    | .attributes[] | select(.name==\"cpuset\").value" | head -1; }

LL_CPUS=$(bln_cpus low-latency)
NORM_CPUS=$(bln_cpus normal)
HK_CPUS=$(bln_cpus default)
echo "low-latency: $LL_CPUS   normal: $NORM_CPUS   housekeeping: $HK_CPUS"

5.2 Allowed CPUs and memories per container

for p in redis-lowlat redis-normal bench housekeeper; do
    printf '%-14s ' "$p"
    kubectl -n demo exec $p -- grep -h -E 'Cpus_allowed_list|Mems_allowed_list' /proc/1/status | tr '\n\t' '  '
    echo
done

The two Redis containers must show disjoint CPU sets of equal size.

5.3 Low-latency CPUs are kernel-isolated, normal CPUs are not

echo "kernel isolcpus: $(cat /sys/devices/system/cpu/isolated)"
echo "low-latency:     $LL_CPUS"
echo "normal:          $NORM_CPUS"

$LL_CPUS must be a subset of isolated; $NORM_CPUS must not intersect it.

Count how many threads in the whole system are allowed to run on each CPU (Cpus_allowed_list in /proc/<pid>/task/<tid>/status):

cat > cpulist.sh <<'EOF'
# expand_cpus "4-7,10" -> "4 5 6 7 10"
expand_cpus() {
    local p
    for p in ${1//,/ }; do
        case $p in *-*) seq "${p%-*}" "${p#*-}" ;; *) echo "$p" ;; esac
    done | tr '\n' ' '
}
EOF

cat > cpu-tasks.sh <<'EOF'
#!/bin/bash
# Usage: sudo ./cpu-tasks.sh CPULIST [-v]
. "$(dirname "$0")/cpulist.sh"
grep -H Cpus_allowed_list /proc/[0-9]*/task/[0-9]*/status 2>/dev/null |
awk -v want="$(expand_cpus "$1")" -v verbose="$2" '
function inlist(c,l,  n,a,j,r){ n=split(l,a,","); for(j=1;j<=n;j++){
    if(split(a[j],r,"-")==2){ if(c>=r[1]+0 && c<=r[2]+0) return 1 } else if(c+0==a[j]+0) return 1 } return 0 }
BEGIN{ m=split(want,w," ") }
{ split($1,p,"/"); pid=p[3]; tid=p[5]; mask=$2
  cmd=""; f="/proc/" pid "/task/" tid "/comm"; if((getline l < f)>0) cmd=l; close(f)
  for(i=1;i<=m;i++) if(inlist(w[i]+0,mask)) { cnt[w[i]]++; seen[w[i] SUBSEP cmd]=1 } }
END{ printf "%-5s %s\n", "CPU", "THREADS-ALLOWED"
  for(i=1;i<=m;i++){ printf "%-5s %d\n", w[i], cnt[w[i]]+0
    if(verbose=="-v") for(k in seen){ split(k,q,SUBSEP); if(q[1]==w[i]) printf "        %s\n", q[2] } } }'
EOF
chmod +x cpu-tasks.sh

sudo ./cpu-tasks.sh "$LL_CPUS"
sudo ./cpu-tasks.sh "$NORM_CPUS"
sudo ./cpu-tasks.sh "$HK_CPUS"

Expected: an order of magnitude fewer threads allowed on the low-latency CPUs. Add -v to list them by name:

sudo ./cpu-tasks.sh "$LL_CPUS" -v

What remains there are per-CPU kernel threads (kworker/<cpu>:*, ksoftirqd/<cpu>, migration/<cpu>, cpuhp/<cpu>), unbound kworker rescuers that only run when a workqueue stalls, and the low-latency pod itself.

This count is the clearest single measure of whether the kernel command line is doing its job. isolcpus=domain alone does not shrink it much: it stops the load balancer from migrating runnable tasks onto the CPU, but unbound kernel threads still list it in their affinity mask and can be woken there. nohz_full is what removes them, because the kernel derives its housekeeping mask for unbound kthreads from it. Drop nohz_full from the command line and re-run this command to see the difference on your node.

5.4 IRQ affinities and interrupt load

cat > irq-load.sh <<'EOF'
#!/bin/bash
# Usage: ./irq-load.sh CPULIST [SECONDS]
. "$(dirname "$0")/cpulist.sh"
cpus=$(expand_cpus "$1"); secs=${2:-5}
affinity_counts() { grep -h . /proc/irq/*/smp_affinity_list 2>/dev/null | awk '
  { n=split($0,p,",")
    for(i=1;i<=n;i++) if(split(p[i],r,"-")==2){for(c=r[1];c<=r[2];c++)cnt[c]++} else cnt[p[i]]++ }
  END{ for(c in cnt) print c, cnt[c] }'; }
delivered() { awk -v c=$1 'NR==1{for(i=1;i<=NF;i++) if($i=="CPU"c) col=i+1; next}
                           col && $col ~ /^[0-9]+$/ {s+=$col} END{print s+0}' /proc/interrupts; }
declare -A aff; while read -r c n; do aff[$c]=$n; done < <(affinity_counts)
declare -A t0; for c in $cpus; do t0[$c]=$(delivered $c); done
sleep "$secs"
printf '%-5s %-14s %s\n' CPU IRQS-AFFINE INTERRUPTS/s
for c in $cpus; do
    printf '%-5s %-14s %s\n' "$c" "${aff[$c]:-0}" \
        "$(( ( $(delivered $c) - ${t0[$c]} ) / secs ))"
done
EOF
chmod +x irq-load.sh

./irq-load.sh "$LL_CPUS,$NORM_CPUS,$HK_CPUS" 5

Expected: the housekeeping CPUs are in the affinity mask of nearly every IRQ and receive nearly all interrupts; the low-latency CPUs are in very few and receive almost none.

The IRQs that remain on the low-latency CPUs are kernel-managed per-CPU device queues (nvme*q*, seen as affinity is read-only warnings in the plugin log). Their affinity cannot be changed from user space. They only fire for I/O submitted from that CPU, which is what isolcpus=managed_irq limits.

Restrict which IRQs the policy touches with controllableInterrupts, for example controllableInterrupts: ["*eth0*", "*eno1*"], to silence those warnings.

5.5 CPU frequency limits, C-states and SST CLOS

cat > cpu-config.sh <<'EOF'
#!/bin/bash
# Usage: sudo ./cpu-config.sh CPULIST
. "$(dirname "$0")/cpulist.sh"
printf '%-5s %-9s %-9s %-12s %-18s %s\n' CPU MIN-MHz MAX-MHz GOVERNOR DISABLED-CSTATES SST-CLOS
for cpu in $(expand_cpus "$1"); do
    d=/sys/devices/system/cpu/cpu$cpu/cpufreq
    dis=$(for s in /sys/devices/system/cpu/cpu$cpu/cpuidle/state*; do
              [ "$(cat $s/disable)" = 1 ] && printf '%s ' "$(cat $s/name)"
          done)
    clos=$(intel-speed-select -c $cpu core-power get-assoc 2>&1 | sed -n 's/.*clos://p')
    printf '%-5s %-9s %-9s %-12s %-18s %s\n' \
        "$cpu" $(( $(cat $d/scaling_min_freq)/1000 )) $(( $(cat $d/scaling_max_freq)/1000 )) \
        "$(cat $d/scaling_governor)" "${dis:--}" "${clos:--}"
done
EOF
chmod +x cpu-config.sh

sudo ./cpu-config.sh "$LL_CPUS,$NORM_CPUS,$HK_CPUS"

Expected: low-latency CPUs at base..turbo MHz, in CLOS 0, with the deep C-states disabled; all other CPUs at min..base MHz, in CLOS 3, all C-states enabled.

SST-CP is enabled and CLOS-based (not proportional), and the two CLOSes carry the frequency windows from pctMinFreq/pctMaxFreq:

sudo intel-speed-select core-power info    2>&1 | head -10
sudo intel-speed-select core-power get-config -c 0 2>&1 | head -11
sudo intel-speed-select core-power get-config -c 3 2>&1 | head -11

The top turbo bucket that SST-TF grants to HP cores, and how many HP cores fit in it per power domain:

LEVEL=$(sudo intel-speed-select perf-profile get-config-current-level 2>&1 \
        | sed -n 's/.*current_level://p' | head -1)
sudo intel-speed-select turbo-freq info -l "$LEVEL" 2>&1 | head -12

Real frequency and C-state residency under load. Run both benchmarks at once and sample the node:

for s in redis-lowlat redis-normal; do
    ( kubectl -n demo exec bench -- memtier_benchmark -s $s -p 6379 \
        --threads=2 --clients=4 --test-time=60 --data-size=32 \
        --key-maximum=10000 --ratio=1:10 --hide-histogram > load-$s.log 2>&1 ) &
done
sleep 6
sudo turbostat --quiet --cpu "$LL_CPUS,$NORM_CPUS" \
    --show CPU,Busy%,Avg_MHz,Bzy_MHz,C1%,C1E%,C6% --interval 8 --num_iterations 3
wait

Bzy_MHz is the frequency while the CPU was busy, C*% the residency in each C-state. Read the rows where Busy% is near 100: each balloon has two CPUs and Redis serves from one thread, so the other CPU of each balloon is idle. If no row has a high Busy%, the sample missed the load — re-run the block.

Expect:

  • low-latency CPU: an HP turbo frequency, matching one of the high-priority-max-level-*-frequency values from the previous command. Which level you get depends on how many cores are active node-wide, so it is not always the bucket-0 value, and it can differ between two runs of the same configuration. Treat sustained-throughput comparisons with that in mind.
  • normal CPU: about the base frequency — not the turbo its cpufreq limit reports. This is the SST-TF LP cap described above: while SST-TF is enabled, the hardware reserves the high turbo buckets for HP cores.
  • 0.00 residency in the C-states listed in disabledCstates, on both the busy and the idle low-latency CPU. The normal balloon's idle CPU shows real C6 residency: that is the wake-up cost section 6.2 measures.

This step only needs to keep the server CPUs busy; ignore any "memtier may be the bottleneck" warning in load-*.log here.

5.6 Scheduling policy and I/O priority

cat > proc-sched.sh <<'EOF'
#!/bin/bash
# Usage: sudo ./proc-sched.sh NAMESPACE POD
uid=$(kubectl -n "$1" get pod "$2" -o jsonpath='{.metadata.uid}')
printf '%-8s %-4s %-5s %-7s %-4s %-22s %s\n' TID CPU CLS RTPRIO NI IO-CLASS/PRIO COMMAND
for pid in $(grep -ls "pod$uid" /proc/[0-9]*/cgroup 2>/dev/null | cut -d/ -f3); do
    while read -r tid cpu cls rt ni comm; do
        [ "$comm" = pause ] && continue
        printf '%-8s %-4s %-5s %-7s %-4s %-22s %s\n' "$tid" "$cpu" "$cls" "$rt" "$ni" \
            "$(ionice -p "$tid" 2>/dev/null | tr -d '\n')" "$comm"
    done < <(ps -Lo tid=,psr=,cls=,rtprio=,ni=,comm= -p "$pid" 2>/dev/null)
done
EOF
chmod +x proc-sched.sh

for p in redis-lowlat redis-normal; do echo "== $p"; sudo ./proc-sched.sh demo $p; done

Expected: every thread of redis-lowlat in class FF (SCHED_FIFO) at RTPRIO 80 and I/O class realtime: prio 0; every thread of redis-normal in class TS (SCHED_OTHER) at NI 0 and I/O class best-effort: prio 7.

CLS/RTPRIO/NI come from ps; the same values are readable per process with chrt -p <tid> and ionice -p <tid>.


6. Measure

Two operating points. The first shows sustained throughput and tail latency, the second shows wake-up latency, where deep C-state exit and frequency ramp-up dominate.

memtier_benchmark reports Ops/sec and the Avg, p50, p99 and p99.9 latency of the whole request/response round trip, including the pod network path that is identical for both servers.

6.1 Sustained load

for s in redis-lowlat redis-normal; do
    echo "=== $s"
    kubectl -n demo exec bench -- memtier_benchmark -s $s -p 6379 \
        --threads=4 --clients=4 --test-time=30 --data-size=32 \
        --key-maximum=10000 --ratio=1:10 --hide-histogram \
        2>&1 | grep -E '^Type|^Totals|Cores used'
done

Cores used shows how many CPUs of the benchmark balloon the driver consumed. If it approaches --threads, raise --threads and, if needed, minCPUs/maxCPUs of the benchmark balloon type, and repeat — otherwise the client, not the server, is being measured. memtier_benchmark prints its own "memtier may be the bottleneck" warning in that case.

6.2 Low request rate — wake-up latency

--rate-limiting=200 sends 200 requests per second per client, so the server CPU is idle between requests and every request pays the wake-up cost.

for s in redis-lowlat redis-normal; do
    echo "=== $s"
    kubectl -n demo exec bench -- memtier_benchmark -s $s -p 6379 \
        --threads=1 --clients=1 --rate-limiting=200 --test-time=30 \
        --data-size=32 --key-maximum=10000 --ratio=1:10 --hide-histogram \
        2>&1 | grep -E '^Type|^Totals'
done

Compare Avg. Latency, p50, p99 and p99.9 between the two servers, and compare C6 exit latency (/sys/devices/system/cpu/cpu0/cpuidle/state*/latency) against the difference you measure.

6.3 Repeat under a containerized background load

In Kubernetes the neighbours are containers, so the background load is a pod. It lands in the noise balloon type: same CPU class and scheduling class as normal, its own dedicated CPUs, sized to fill the rest of $DEMO_NODE. Neither Redis instance loses CPU time to it — both keep their dedicated CPUs. What the noise does take away is package-level turbo budget, last-level cache and memory bandwidth, which is exactly the contention pctPriority arbitrates.

kubectl apply -f noise.yaml
kubectl -n demo wait --for=condition=Ready pod/noise --timeout=120s
kubectl get noderesourcetopologies -o json | jq -r '
  .items[].zones[] | select(.type=="balloon")
  | [.name, ((.attributes[]|select(.name=="cpuset").value) // "-")] | @tsv' | column -t -s $'\t'

The noise[0] cpuset must be disjoint from low-latency[0], normal[0] and benchmark[0], and inside $DEMO_NODE.

Frequency under contention — this is where the HP and LP CLOSes separate:

NOISE_CPUS_ACTUAL=$(kubectl get noderesourcetopologies -o json | jq -r \
  '.items[].zones[]|select(.name|startswith("noise"))|.attributes[]|select(.name=="cpuset").value')
sudo turbostat --quiet --cpu "$LL_CPUS,$NORM_CPUS,$NOISE_CPUS_ACTUAL" \
    --show CPU,Busy%,Bzy_MHz --interval 8 --num_iterations 2

Repeat both measurements:

for s in redis-lowlat redis-normal; do
    echo "=== $s, sustained, with noise"
    kubectl -n demo exec bench -- memtier_benchmark -s $s -p 6379 \
        --threads=4 --clients=4 --test-time=30 --data-size=32 \
        --key-maximum=10000 --ratio=1:10 --hide-histogram \
        2>&1 | grep -E '^Type|^Totals|Cores used'
done

for s in redis-lowlat redis-normal; do
    echo "=== $s, 200 req/s, with noise"
    kubectl -n demo exec bench -- memtier_benchmark -s $s -p 6379 \
        --threads=1 --clients=1 --rate-limiting=200 --test-time=30 \
        --data-size=32 --key-maximum=10000 --ratio=1:10 --hide-histogram \
        2>&1 | grep -E '^Type|^Totals'
done

Remove the load again:

kubectl delete -f noise.yaml

Two effects to separate when reading the numbers:

  • The low-latency CPUs sit in the HP CLOS with pctMinFreq: turbo, so their frequency should hold while the LP cores around them lose turbo headroom.
  • A busy neighbour also keeps its own CPU out of deep C-states. It cannot do that for the normal Redis CPU, because that CPU is dedicated and still idles between requests — so the low request rate gap does not close.

6.4 What the kernel command line and the governor are worth

Three configuration choices from section 1.1 and section 3 that only measurement can settle. Run 6.1 and 6.2 again after each change and compare.

irqaffinity — add irqaffinity=<every CPU except $ISOLCPUS> to the kernel command line, reboot, and re-run section 5.4 rather than the benchmarks: the question is whether the static boot mask reaches any interrupt that irqMode does not.

./irq-load.sh "$LL_CPUS,$NORM_CPUS,$HK_CPUS" 5

Compare both columns. A lower IRQS-AFFINE on the low-latency CPUs with no change in INTERRUPTS/s means the static mask only narrowed interrupts that never fire.

nohz_full — reboot with nohz_full=$ISOLCPUS removed from the kernel command line, keeping everything else identical:

cat /sys/devices/system/cpu/nohz_full     # the isolated CPUs, or empty
sudo ./cpu-tasks.sh "$LL_CPUS"            # compare against section 5.3

Two effects pull in opposite directions, so measure both columns:

  • A nohz_full CPU skips the periodic tick while a single task runs, but pays extra tick and RCU bookkeeping on every user/kernel transition. Redis crosses that boundary several times per request, so tick suppression is not obviously a win for it.
  • Without nohz_full the kernel has no housekeeping mask for unbound kernel threads, and the number of threads allowed on the low-latency CPUs grows by roughly a factor of three. Those threads rarely run, but nothing stops them.

freqGovernor — the low-latency class already pins minFreq to turbo. Test whether the governor still matters on top of that by switching the low-latency class between performance and powersave:

llclass() {   # patch freqGovernor of the low-latency cpuClass in place
    local idx
    idx=$(kubectl -n kube-system get balloonspolicy default -o json \
          | jq '[.spec.cpuClasses[].name] | index("low-latency")')
    kubectl -n kube-system patch balloonspolicy default --type=json \
      -p "[{\"op\":\"replace\",\"path\":\"/spec/cpuClasses/$idx/freqGovernor\",\"value\":\"$1\"}]"
}

llclass powersave
sleep 10
sudo ./cpu-config.sh "$LL_CPUS"          # governor changed, min and max still turbo?
# repeat 6.1 and 6.2 here

llclass performance
sleep 10
sudo ./cpu-config.sh "$LL_CPUS"

The jq index lookup keeps the patch on the low-latency class only; editing the YAML with sed would also hit the default class, which must stay on powersave.

6.5 How much of the difference is PCT?

With SST-TF enabled the hardware caps LP cores near base frequency, so part of the sustained-throughput gap is the CLOS priority rather than anything else in the configuration. Separate the two by taking PCT out entirely and repeating 6.1 and 6.2:

sed '/^ *pct\(Priority\|MinFreq\|MaxFreq\):/d' low-latency-balloons.yaml \
    > no-pct-balloons.yaml
diff low-latency-balloons.yaml no-pct-balloons.yaml
kubectl replace -f no-pct-balloons.yaml
sleep 12

# the policy does not undo SoC-wide SST state when PCT disappears from the config
sudo intel-speed-select core-power disable
sudo intel-speed-select turbo-freq disable
sudo intel-speed-select core-power info 2>&1 | head -9      # enable-status:disabled

sudo ./cpu-config.sh "$LL_CPUS,$NORM_CPUS"    # cpufreq limits unchanged
# repeat 6.1 and 6.2 here

Now both classes really do have the same frequency ceiling, and the remaining low-latency advantage comes only from the pinned frequency floor, the disabled C-states, isolcpus, the IRQ isolation and the realtime scheduling policy.

The SST-CLOS column of cpu-config.sh keeps reporting the last association even after core-power disable; core-power info is the field that tells you whether CLOSes are in effect at all.

Put PCT back afterwards:

kubectl replace -f low-latency-balloons.yaml
sleep 12
sudo intel-speed-select core-power info 2>&1 | head -9      # enable-status:enabled
sudo ./cpu-config.sh "$LL_CPUS,$NORM_CPUS"                  # CLOS 0 and 3 again

7. Reset the node and uninstall

Order matters. Delete the workloads first: scheduling policy and I/O priority are set at container creation and are not reverted by a configuration change.

kubectl delete -f noise.yaml --ignore-not-found
kubectl delete -f workloads.yaml

Apply a reset policy that puts every container into one balloon covering all CPUs, restores cpufreq limits and C-states, spreads IRQs back over all CPUs, and returns SST to CLOS 0 with the full frequency range. This mirrors Reset CPU and Memory Pinning.

N=$(nproc --all)
cat > reset-balloons.yaml <<EOF
apiVersion: config.nri/v1alpha1
kind: BalloonsPolicy
metadata:
  name: default
  namespace: kube-system
spec:
  pinCPU: true
  pinMemory: true
  reservedResources:
    cpu: cpuset:0-$((N-1))
  balloonTypes:
  - name: reserved
    namespaces: ["*"]        # every container into this balloon
    minCPUs: $N              # balloon covers every CPU
    maxCPUs: $N
    preferIsolCpus: true     # including the kernel-isolated ones
    irqClaim: ["*"]          # every IRQ back onto every CPU
  cpuClasses:
  # Applied to balloon CPUs and to idle CPUs: full frequency range, the
  # governor the node had before the demo, no C-state disabled, all CPUs
  # into SST CLOS 0 with no cap.
  - name: default
    minFreq: min
    maxFreq: turbo
    freqGovernor: $ORIG_GOV
    pctPriority: high
    pctMinFreq: min
    pctMaxFreq: turbo
EOF

kubectl replace -f reset-balloons.yaml
sleep 20

Verify the reset:

sudo ./cpu-config.sh 0,1,2,$((N-1))          # min..turbo, no disabled C-state, CLOS 0
./irq-load.sh 0,1,2,$((N-1)) 5               # every CPU in nearly every IRQ affinity
sudo intel-speed-select core-power get-config -c 0 2>&1 | head -11
kubectl get noderesourcetopologies -o json | jq -r \
  '.items[].zones[]|select(.type=="balloon")|[.name,((.attributes[]|select(.name=="cpuset").value)//"-")]|@tsv'

# cpusets of every running container: only "0-<N-1>" (and empty) must be left
find /sys/fs/cgroup/kubepods* -name cpuset.cpus -print0 | xargs -0 cat | sort -u

Uninstall. The plugin does not turn SST off on exit, so disable SST-CP and SST-TF explicitly — this restores enable-status:disabled, clos-enable-status:disabled, priority-type:proportional. See Uninstalling with Helm.

helm uninstall nri-resource-policy-balloons -n kube-system
kubectl delete crd balloonspolicies.config.nri

sudo intel-speed-select core-power disable
sudo intel-speed-select turbo-freq disable
sudo intel-speed-select core-power info 2>&1 | head -10

Final check that nothing is left tuned:

grep -h . /sys/devices/system/cpu/cpu*/cpufreq/scaling_min_freq   | sort -u
grep -h . /sys/devices/system/cpu/cpu*/cpufreq/scaling_max_freq   | sort -u
grep -h . /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor   | sort -u   # $ORIG_GOV
grep -l 1 /sys/devices/system/cpu/cpu*/cpuidle/state*/disable | wc -l         # 0

If the last count is not 0, the plugin pod was restarted between disabling the C-states and applying the reset policy: its cpuidle writer is created lazily, so a fresh process with no disabledCstates anywhere never touches cpuidle sysfs. Re-enable them directly:

for f in $(grep -l 1 /sys/devices/system/cpu/cpu*/cpuidle/state*/disable); do
    echo 0 | sudo tee "$f" > /dev/null
done
grep -l 1 /sys/devices/system/cpu/cpu*/cpuidle/state*/disable | wc -l   # 0

Two things are not restored by these commands:

  • The exact boot-time IRQ affinity from the irqaffinity= kernel parameter. The reset policy spreads IRQs over all CPUs, including the isolated ones. A reboot restores the boot-time masks.
  • The isolcpus, nohz_full and rcu_nocbs kernel parameters. Remove them from /etc/default/grub (or restore /etc/default/grub.bak), run sudo update-grub and reboot.
sudo cp /etc/default/grub.bak /etc/default/grub && sudo update-grub && sudo reboot

Reference

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors