← All Advisories

Linux Kernel TC Classifier Filter Allocations in the change() Path Use Plain GFP Flags Without Accounting to memcg, Allowing Container Workloads to Exhaust Host Memory Without Being Charged

Last refreshed2026-09-28

Status: NEW  |  Advisory ID: CVE-2026-90099

Key Details

CVECVE-2026-90099

What to Know

In the Linux kernel, the following vulnerability has been resolved:

net/sched: account classifier filter allocations to memcg

Allocations in the tc classifier *_change() paths (filter objects,

per-CPU counters, and per-filter aux data) use plain GFP_KERNEL without

__GFP_ACCOUNT, allowing unprivileged users to pin kernel memory outside

memcg charging. The shared tcf_exts_init_ex() action array allocation in

cls_api.c was also uncharged; this patch closes it along with the

per-classifier filter-object/percpu/aux allocations that remain

unaccounted.

Add GFP_KERNEL_ACCOUNT to:

- the shared tcf_exts_init_ex() action array (cls_api.c), common to every

filter of every classifier (32 pointers, 256 bytes);

- the filter-object, per-CPU-counter, and per-filter aux allocations in

cls_basic, cls_bpf, cls_cgroup, cls_flow, cls_flower, cls_fw,

cls_matchall, cls_route and cls_u32;

- the u32_init_knode() replace-path knode allocation (cls_u32.c), which

allocates the same struct tc_u_knode + sel.keys on every replace of an

existing knode and was missed by the create-path-only conversion.

Also fix the cls_basic error path: basic_change() inserts fnew into the

IDR before allocating the per-CPU counter. If alloc_percpu() fails the

errout path kfree'd fnew without idr_remove, leaving a dangling pointer

in the IDR. With GFP_KERNEL_ACCOUNT the percpu alloc becomes failable

on demand (memcg at memory.max), making the dead path attacker-reachable

and burning the handle permanently. Add the idr_remove on the percpu

failure path, matching the basic_set_parms failure-path pattern.

Note: [email protected] provided a poc for basic_cls, but it was easy to

extend to the other classifiers.

Conditions to recreate the bug:

- CONFIG_NET_SCHED, CONFIG_NET_CLS_* (the classifier being used),

CONFIG_NET_CLS_ACT, CONFIG_MEMCG, CONFIG_USER_NS, CONFIG_NET_NS.

- Unprivileged user in a fresh user+network namespace (unshare -Urn),

or root with CAP_NET_ADMIN.

- Create a large number of tc filters (e.g. tc filter add dev lo

ingress ... <classifier> ...) while watching a memcg-limited cgroup:

system slab grows far faster than memory.current, pinning kernel

memory outside memcg charging. (NVD)

References

SourceReference
NVDhttps://nvd.nist.gov/vuln/detail/CVE-2026-90099
CVEhttps://www.cve.org/CVERecord?id=CVE-2026-90099