Merge tag 'mm-hotfixes-stable-2026-09-27-19-12' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Pull MM fixes from Andrew Morton:
- Fix module loading incorrectly returning success after alloc_tag
codetag setup failed
- Restore MADV_COLLAPSE semantics for shmem so forced collapse ignores
the shmem THP/mTHP sysfs settings
- Make alloc_tag UAPI structure padding explicit
- Fix DAMON schemes unexpectedly stopping after quotas are disabled
through the online parameter update interface
- Fix a false SW_TAGS KASAN invalid-access report when freeing vmapped
task stacks
- Split the MEMORY MANAGEMENT - MEMORY POLICY AND MIGRATION MAINTAINERS
[22 lines not shown]
Merge tag 'driver-core-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/driver-core/driver-core
Pull driver core fix from Danilo Krummrich:
- Suppress spurious "debugfs is not initialized yet" boot warnings when
the caller passes an error parent to debugfs file creation; callers
propagating an earlier failure should not trigger the warning
* tag 'driver-core-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/driver-core/driver-core:
debugfs: don't warn about uninitialized debugfs for an error parent
Merge tag 'i2c-fixes-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/andi.shyti/linux
Pull i2c fix from Andi Shyti:
- qcom-geni: select the correct source clock table entry
* tag 'i2c-fixes-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/andi.shyti/linux:
i2c: qcom-geni: Fix hardcoded clock index in SE_GENI_CLK_SEL
Merge tag 'wq-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/wq
Pull workqueue fix from Tejun Heo:
- Fix a NULL dereference in the flush dependency check when a worker
flushes outside a work item, such as from the OOM path during worker
creation
* tag 'wq-for-7.3-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/wq:
workqueue: Fix NULL current_pwq deref in flush dependency check
Merge tag 'cgroup-for-7.3-rc4-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup
Pull cgroup fix from Tejun Heo:
- A cpuset partition could claim CPUs an ancestor partition already
held exclusively. Restore the rejection an earlier change had turned
into a warning.
* tag 'cgroup-for-7.3-rc4-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup:
cgroup/cpuset: Return PERR_NOCPUS in remote_partition_enable() on subpartitions_cpus conflict
Merge tag 'sched_ext-for-7.3-rc4-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext
Pull sched_ext fix from Tejun Heo:
- The CPU topology helper for BPF schedulers took no buffer size, so
its structure couldn't grow without breaking schedulers built against
the older layout. Add a size argument.
* tag 'sched_ext-for-7.3-rc4-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext:
sched_ext: Add a size argument to scx_bpf_cid_topo() so struct scx_cid_topo can grow
Merge tag 'x86-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull x86 fixes from Ingo Molnar:
- Fix preemption bugs in the SVSM vTPM guest implementation
(Melody Wang)
- Fix MCE-triggered hardware debug register corruption on
task migration (Masami Hiramatsu)
* tag 'x86-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
x86/mce: Fix hardware debug register corruption on task migration
x86/sev: Make vTPM SVSM calls preemption-safe
Merge tag 'sched-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull scheduler fixes from Ingo Molnar:
- Fix LLC mis-scheduling bugs (Tim Chen, Lu Wang)
- Fix cache-grouping related scheduling statistics UAF bugs (Tim Chen)
- Skip kernel threads for cache aware scheduling to rubustify the code
(Chen Yu)
- Refresh LLC capacity across CPU hotplug, to fix capacity
underestimation bug (Davi Chaves Azevedo)
- Account PSI IRQ time to the execution context, not the scheduling
context, to fix proxy scheduling accounting bug (Zhan Xusheng)
* tag 'sched-urgent-2026-09-27' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
sched/core: Account PSI IRQ time to the execution context, not the scheduling context
[6 lines not shown]
workqueue: Fix NULL current_pwq deref in flush dependency check
check_flush_dependency() uses current_wq_worker() to determine whether
the caller is a workqueue worker and then dereferences worker->current_pwq
to test whether the current workqueue is WQ_MEM_RECLAIM.
current_wq_worker() only means that %current has PF_WQ_WORKER set. A
kworker can reach check_flush_dependency() while it is not executing a
work item. One such path is worker_thread() acting as the pool manager,
where create_worker() does GFP_KERNEL allocation and the allocation path
invokes the OOM notifier. In that state worker->current_pwq is NULL
because current_pwq is set only by process_one_work() and cleared again
after the work function returns.
[ 416.760634][ T375] Call trace:
[ 416.760638][ T375] check_flush_dependency+0x80/0x120 (P)
[ 416.760648][ T375] __flush_work+0x98/0x224
[ 416.760657][ T375] flush_work+0x30/0x44
[ 416.760665][ T375] ...
[26 lines not shown]
mm/vma: predicate setting mmap_prepare VMA fields on new vma alloc
It only makes sense to manipulate VMA fields if a new VMA was allocated,
rather than merged.
VMA merging does not compare vm_ops or vm_private_data, so a merged VMA
keeps its own, which is also what the legacy f_op->mmap path does since it
never touches an existing VMA.
Currently, these fields will get overwritten by whatever state is
established in the mmap_prepare hook, and if the VMA was merged,
vm_ops->mapped will not have been called, so this could destructively
clear existing state without replacing it with anything valid.
There is an implicit requirement that vm_private_data and vm_ops are
fungible across VMAs which means that losing the 'new' state is fine.
However in this case the 'old' state is being overwritten by potentially
invalid 'new' state, so this must be rectified.
[18 lines not shown]
MAINTAINERS: add Heming Zhao as ocfs2 reviewer
Heming Zhao has contributed ocfs2 a lot recent years, both as a author and
reviewer. So add him as ocfs2 reviewer.
Link: https://lore.kernel.org/20260923003920.2382398-1-joseph.qi@linux.alibaba.com
Signed-off-by: Joseph Qi <joseph.qi at linux.alibaba.com>
Signed-off-by: Andrew Morton <akpm at linux-foundation.org>
Acked-by: Mark Fasheh <mark at fasheh.com>
Cc: Heming Zhao <heming.zhao at suse.com>
Cc: Joel Becker <jlbec at evilplan.org>
MAINTAINERS: make Gregory a co-maintainer of MEMORY MANAGEMENT - NUMA PLACEMENT
Gregory is extremely familiar with NUMA placement handling and volunteer
to help co-maintain the numa placement bits. So add him as a
co-maintainer.
Link: https://lore.kernel.org/20260918-maintainers-mempolicy-v1-3-9a94cac6d135@kernel.org
Signed-off-by: David Hildenbrand (Arm) <david at kernel.org>
Signed-off-by: Andrew Morton <akpm at linux-foundation.org>
Acked-by: Lorenzo Stoakes (ARM) <ljs at kernel.org>
Acked-by: Gregory Price (Meta) <gourry at gourry.net>
Acked-by: Zi Yan <ziy at nvidia.com>
Reviewed-by: Joshua Hahn <joshua.hahnjy at gmail.com>
Acked-by: SJ Park <sj at kernel.org>
Acked-by: Vlastimil Babka (SUSE) <vbabka at kernel.org>
Cc: Alistair Popple <apopple at nvidia.com>
Cc: Byungchul Park <byungchul at sk.com>
Cc: "Huang, Ying" <ying.huang at linux.alibaba.com>
Cc: Liam R. Howlett <liam at infradead.org>
[5 lines not shown]
MAINTAINERS: move memory tiering under MEMORY MANAGEMENT - NUMA PLACEMENT
NUMA PLACEMENT is a better place for memory tiering.
My best guess is that existing MISC reviewers are not that interested in
memory tiering, so don't carry any over.
Link: https://lore.kernel.org/20260918-maintainers-mempolicy-v1-2-9a94cac6d135@kernel.org
Signed-off-by: David Hildenbrand (Arm) <david at kernel.org>
Signed-off-by: Andrew Morton <akpm at linux-foundation.org>
Acked-by: Lorenzo Stoakes (ARM) <ljs at kernel.org>
Acked-by: Zi Yan <ziy at nvidia.com>
Reviewed-by: Joshua Hahn <joshua.hahnjy at gmail.com>
Reviewed-byt: SJ Park <sj at kernel.org>
Acked-by: Vlastimil Babka (SUSE) <vbabka at kernel.org>
Cc: Alistair Popple <apopple at nvidia.com>
Cc: Byungchul Park <byungchul at sk.com>
Cc: Gregory Price <gourry at gourry.net>
Cc: "Huang, Ying" <ying.huang at linux.alibaba.com>
[6 lines not shown]
MAINTAINERS: split up MEMORY MANAGEMENT - MEMORY POLICY AND MIGRATION
Patch series "MAINTAINERS: rework MEMORY MANAGEMENT - MEMORY POLICY".
Let's rework MEMORY MANAGEMENT - MEMORY POLICY AND MIGRATION.
As I can use some maintenance help in that area, add Gregory as a new
co-maintainer for the split out MEMORY MANAGEMENT - NUMA PLACEMENT
section.
This patch (of 3):
Let's split it up into MIGRATION and NUMA PLACEMENT. The latter is a
better fitting description for mempolicy.c.
Make a best guess about which pieces existing reviewers are interested in.
Move the split sections to keep alphabetical order.
[18 lines not shown]
mm/damon/core: don't skip damos_adjust_quota() while esz is not zero
DAMOS could unexpectedly stop working when a user disables quota using the
online parameters commit feature. Fix it by correcting a wrong quota
unset check in damos_adjust_quota().
DAMON users could disable all quotas by unsetting time and size quotas,
and removing all quota goals. The intention of disabling quotas would be
making DAMOS run at full speed. When such quota disabled setup is
detected, damos_adjust_quota() skips all its work. The skipped works
include effective size quota (damos_quota->esz) updates and charged quota
amount (damos_quota->charged_sz) resets. The intention is to avoid doing
unnecessary work when quotas are disabled.
However, users could do the setup while effective size quota is non-zero,
by doing the disabling with the online DAMON parameters commit feature.
In this case, because the effective size quota exists, DAMOS will keep
working with the quota until it is fully charged. After the effective
quota is fully charged, the charged quota amount (damos_quota->charged_sz)
[30 lines not shown]
alloc_tag: avoid implicit padding in uapi
The implied padding causes a harmless warning when testing the uapi
headers with -Wpadded that could in theory indicate incompatibilities
or data leaks:
./usr/include/linux/alloc_tag.h:41:1: error: padding struct size to alignment boundary with 7 bytes [-Werror=padded]
The code here is fine, but it's better to make the padding explicit and
avoid the warning here.
Link: https://lore.kernel.org/20260916065830.1619425-1-arnd@kernel.org
Link: https://lore.kernel.org/20260915202404.3568029-1-arnd@kernel.org
Fixes: 1d581ab2348c ("alloc_tag: add ioctl to /proc/allocinfo")
Signed-off-by: Arnd Bergmann <arnd at arndb.de>
Signed-off-by: Andrew Morton <akpm at linux-foundation.org>
Acked-by: Suren Baghdasaryan <surenb at google.com>
Acked-by: SJ Park <sj at kernel.org>
Acked-by: Hao Ge <hao.ge at linux.dev>
kasan: unpoison task stack below watermark only in generic mode
CPU resume and BPF exception handling can discard stack frames without
running their compiler-generated epilogues. Generic KASAN needs
kasan_unpoison_task_stack_below() to clear the redzones left behind by
those frames before the stack is reused.
CONFIG_KASAN_STACK also enables this helper for SW_TAGS. The helper
derives the stack base from an untagged stack pointer, so kasan_unpoison()
writes KASAN_TAG_KERNEL (0xff) into the shadow for [base, watermark). For
a vmapped task stack, vm_area->addr still carries the original random
allocation tag. This creates a tag mismatch at the stack base even if
that memory has never held an instrumented stack object.
When the task exits and its stack is not cached, thread_stack_free_rcu()
passes vm_area->addr to vfree(). In RCU callback context this reaches
vfree_atomic(), whose llist_add() writes to the allocation base through
the tagged pointer and triggers a false KASAN invalid-access report:
[38 lines not shown]
mm: shmem: ignore sysfs configs for shmem forced collapse
According to Documentation/mm/transhuge.rst, MADV_COLLAPSE is expected to
ignore any THP or mTHP interface settings.
However, after commit 26c7d8413aaf ("mm: thp: support "THPeligible"
semantics for mTHP with anonymous shmem"), performing MADV_COLLAPSE on
shmem will depend on /sys/.../hugepages-2048kB/shmem_enabled being set to
"inherit" (although that is the default), which can cause MADV_COLLAPSE to
fail unexpectedly.
This could cause a userspace-visible performance regression.
Fix this by returning the result of shmem_huge_global_enabled() directly
when MADV_COLLAPSE is requested, allowing PMD-order collapse while
ignoring shmem THP/mTHP settings.
Link: https://lore.kernel.org/063f655b4d6c4234f3aa27ed6ecab10283ecb880.1789351825.git.baolin.wang@linux.alibaba.com
Fixes: 26c7d8413aaf ("mm: thp: support "THPeligible" semantics for mTHP with anonymous shmem")
[16 lines not shown]
module: fix lost error code from codetag_load_module()
If codetag_load_module() fails, err is not set to reflect the failure
and load_module() returns 0 after the module has been torn down.
Also, if the module is a livepatch, mod->klp_info allocated by
copy_module_elf() leaks on this error path. Free it via a new
livepatch_cleanup label.
Link: https://lore.kernel.org/20260827030503.49171-1-hao.ge@linux.dev
Fixes: 044d2aee6c57 ("alloc_tag: handle module codetag load errors as module load failures")
Signed-off-by: Hao Ge <hao.ge at linux.dev>
Signed-off-by: Andrew Morton <akpm at linux-foundation.org>
Reported-by: Sashiko <sashiko-bot at kernel.org>
Suggested-by: Petr Pavlu <petr.pavlu at suse.com>
Reviewed-by: Bradley Morgan <brads at mainlining.org>
Cc: Aaron Tomlin <atomlin at atomlin.com>
Cc: Luis Chamberalin <mcgrof at kernel.org>
Cc: Sami Tolvanen <samitolvanen at google.com>
[2 lines not shown]
sched_ext: Add a size argument to scx_bpf_cid_topo() so struct scx_cid_topo can grow
scx_bpf_cid_topo() copies struct scx_cid_topo into a buffer the BPF program
sized from its own vmlinux.h while the verifier sizes the write from the
running kernel's BTF. The struct may grow and each growth then breaks every
scheduler built against the older layout, rejected at load or written past
its buffer. This is the usual hole for a struct handed to BPF, closed
elsewhere with a size argument, and it was missed here.
Take the buffer size, copy the smaller of it and the kernel's struct and set
the rest to -1. Accesses to the copy are CO-RE relocated, so the struct can
grow by appending fields, which its comment now states. The kfunc changes in
place: the cid interface is still being finalized and no released scheduler
uses the current form.
Fixes: e9b55af47edf ("sched_ext: Add topological CPU IDs (cids)")
Cc: stable at vger.kernel.org # v7.2+
Signed-off-by: Tejun Heo <tj at kernel.org>
Reviewed-by: Andrea Righi <arighi at nvidia.com>
Merge tag 'ata-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/libata/linux
Pull ata fixes from Niklas Cassel:
- Extend the quirk "no LPM on ATI" quirk, that is currently only
applied for Samsung drives, to include AMD controllers as well.
The AMD AHCI controllers are newer versions of the ATI AHCI
controllers, and these controllers still have LPM issues with
Samsung drives - LPM works with drives from other vendors (me)
- Fix errors in the libata.force parameter documentation (me)
- Verify the sense data descriptor lengths for ATA PASS-THROUGH
command, so that a malicious device cannot write past the buffer
length (Matthias)
- Mention the libata for-next branch in MAINTAINERS such that the
git ls-remote command done by get_maintainer.pl --self-test=scm
[7 lines not shown]
cgroup/cpuset: Return PERR_NOCPUS in remote_partition_enable() on subpartitions_cpus conflict
When a remote partition is created underneath an existing local partition
via a non-partition (PRS_MEMBER) intermediate cgroup, update_prstate() sees
parent->partition_root_state == PRS_MEMBER and calls
remote_partition_enable().
Commit 86888c7bd117 ("cgroup/cpuset: Add warnings to catch inconsistency
in exclusive CPUs") replaced the cpumask_intersects(tmp->new_cpus,
subpartitions_cpus) error check in remote_partition_enable() with
WARN_ON_ONCE(). As a result, remote_partition_enable() emits a warning
and proceeds to enable the remote partition on CPUs that are already
owned by the ancestor local partition in subpartitions_cpus.
This can be reproduced on Linux 7.3.0-rc3 with:
mkdir -p /tmp/cg1
mount -t cgroup2 none /tmp/cg1
echo "+cpuset" > /tmp/cg1/cgroup.subtree_control
[40 lines not shown]
Merge tag 'pci-v7.3-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci
Pull PCI fixes from Bjorn Helgaas:
- Make BAR resize work even for devices where no upstream bridge is
visible to the OS, which fixes an amdgpu regression on SolidRun
HoneyComb, which doesn't expose Root Ports to the OS (Liz Fong-Jones)
- Omit bus properties in dynamic OF nodes when a bridge has no
subordinate bus, which fixes early boot hangs caused by NULL pointer
dereferences with CONFIG_PCI_DYNAMIC_OF_NODES enabled (Angel J)
- Disable enhanced atomics on AMD NBIO 7.7 and 7.11 to avoid silent
data corruption on 64-bit DMAs (Mario Limonciello)
* tag 'pci-v7.3-fixes-2' of git://git.kernel.org/pub/scm/linux/kernel/git/pci/pci:
x86/PCI: Disable enhanced atomics on AMD NBIO 7.7 and 7.11
PCI: of_property: Omit bus properties without a subordinate bus
PCI: Fix BAR resize for devices on a root bus
Merge tag 'probes-fixes-v7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace
Pull probe fixes from Masami Hiramatsu:
- kprobes: Fix permanent hang when flushing the kprobe optimizer
Fix a deadlock when disabling kprobe optimization via sysctl or
debugfs where flushers hung waiting for optimizer_completion.
Replaced the completion with an optimizer_passes counter and
wait_var_event_mutex() under kprobe_mutex so concurrent flushers can
wait and wake up safely.
- fprobe: Terminate the fgraph_data list when the reservation is not
filled
Fix an issue where unused shadow stack data left uninitialized by
fprobe_fgraph_entry() was misparsed as stale fprobe headers on
return. Explicitly write a zero word to terminate the list and update
read_fprobe_header() to handle the zeroed slot properly.
[12 lines not shown]
Merge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm
Pull kvm fixes from Paolo Bonzini:
"Arm:
- Invalidate the ITS translation cache when the guest changes the
base address of the ITS tables (Fuad Tabba)
- Skip saving ITS devices with device IDs that are out-of-bounds
rather than failing the entire ITS save ioctl (Fuad Tabba)
- Close race between VM teardown and invalidations of nested MMUs
when handling MMU operations that are allowed to block (Lorenzo
Stoakes)
- Various fixes for the handling of the host's untrusted SVE
configuration in pKVM (Fuad Tabba)
- Make sure that empty SMCCC ranges based at 0 are rejected by the
[131 lines not shown]
Merge tag 'kvm-x86-fixes-7.3-rc5' of https://github.com/kvm-x86/linux into HEAD
KVM fixes for 7.3-rcN
- Fix a brown paper bag bug where KVM would incorrectly treat Intel PMU MSRs
as valid on AMD.
- Fix a regression in the hardware disable selftest where it checked the wrong
macro when detecting glibc support (breaks at least musl).
- Never clear KVM_REQ_VM_DEAD so that dead VMs stay dead, which is especially
important for KVM_BUG_ON() flows, which often guard more dangerous bugs.
- Re-pend GET_NESTED_STATE_PAGES if getting the pages fails, to fix a bug
where KVM would let userspace run a broken setup with stale vmcs12 pages.
- Fix a class of bugs where KVM would fail to fill kvm_run exit fields if
getting nested pages failed.
[10 lines not shown]
KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV
Unconditionally free SEV's "have run CPUs" cpumask in the VM destroy path,
i.e. even for what appear to be non-SEV VMs, as an SEV VM becomes a non-SEV
VM if its state is intra-host migrated. Alternatively, the mask could be
freed in sev_migrate_from() when "converting" the source VM, but that gets
annoying because ideally KVM would nullify the mask to guard against UAF,
and nullifying the mask would need be conditioned on CPUMASK_OFFSTACK=y.
Freeing the mask during sev_migrate_from() is also not robust against other
KVM bugs, though that's kind of a moot point since any such bugs would show
up even if sev->active is never set. I.e. KVM must get that side of things
correct. But, that's not a great reason to add more code just to make
things marginally less robust.
Fixes: 6f38f8c57464 ("KVM: SVM: Flush cache only on CPUs running SEV guest")
Cc: stable at vger.kernel.org
Reported-by: Stefan Teodorescu <fane at google.com>
Signed-off-by: Sean Christopherson <seanjc at google.com>
[2 lines not shown]
KVM: SEV: Do cache maintenance on the source VM during intra-host migration
Manually perform cache maintenance on the source VM during intra-host
migration to ensure no stale data is left in CPU caches after the VM is
destroyed. Because the source VM is "converted" to a non-SEV VM, KVM's
memory reclaim flows won't trigger cache maintenance, e.g. when all guest
memory is reclaimed in response to detaching from the mmu_notifier.
Note, relying on the destination VM to do cache maintenance isn't an option
as KVM doesn't require identical guest memory configurations, i.e. the
source VM may have access to memory that the destination VM does not.
Enforcing equivalent memory configurations is infeasible, as it would
require a *deep* comparison of memslots, e.g. to verify that not only are
the memslot identical, but what the memslots point at is also identical.
Fixes: b56639318bb2 ("KVM: SEV: Add support for SEV intra host migration")
Cc: stable at vger.kernel.org
Reported-by: Stefan Teodorescu <fane at google.com>
Signed-off-by: Sean Christopherson <seanjc at google.com>
[2 lines not shown]