Changelog in Linux kernel 7.2.9

 
accel/ivpu: Drop IRQF_ONESHOT to allow IPC IRQ threading on PREEMPT_RT [+ + +]
Author: Karol Wachowski <[email protected]>
Date:   Wed Jun 17 11:20:31 2026 +0200

    accel/ivpu: Drop IRQF_ONESHOT to allow IPC IRQ threading on PREEMPT_RT
    
    commit 799c8f0b9f3fd728b32594aed3852044703b7e5b upstream.
    
    The IPC RX hardirq handler matches consumers under a spinlock and
    allocates rx_msg buffers. On PREEMPT_RT these spinlocks become sleeping
    locks and the allocation may sleep, neither of which is allowed in true
    hardirq context, resulting in "sleeping function called from invalid
    context" splats.
    
    IRQF_ONESHOT makes genirq keep the primary handler in hardirq even when
    forced threading is enabled, so on PREEMPT_RT the handler cannot be
    threaded. Drop the flag so the primary handler is threaded on PREEMPT_RT
    and the IPC RX path runs in a context where sleeping is allowed. On the
    MSI interrupt chip (IRQCHIP_ONESHOT_SAFE) the flag was stripped anyway,
    so non-RT behaviour is unchanged.
    
    Fixes: 85c9cc2d25f8 ("accel/ivpu: Use threaded IRQ for IPC callback processing")
    Cc: Andrzej Kacprowski <[email protected]>
    Cc: Karol Wachowski <[email protected]>
    Cc: [email protected]
    Reviewed-by: Andrzej Kacprowski <[email protected]>
    Signed-off-by: Karol Wachowski <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

accel/ivpu: Use separate flag for job timeout [+ + +]
Author: Jakub Pawlak <[email protected]>
Date:   Tue Sep 29 08:07:25 2026 -0400

    accel/ivpu: Use separate flag for job timeout
    
    [ Upstream commit 72b782097e534ab48152dd25aaba40dda5f25f9e ]
    
    Use separate flag to mark a job timeout as a reason
    of starting context_abort_work. This allows to distinguish
    engine reset reason and clearly adjust reset procedure flow.
    
    The flag is cleared in ivpu_prepare_for_reset(), which every
    recovery and suspend path already funnels through, so that the
    state is clean after recovery.
    
    Cc: [email protected] # v7.1+
    Fixes: ade00a6c903f ("accel/ivpu: Perform engine reset instead of device recovery on TDR")
    Signed-off-by: Jakub Pawlak <[email protected]>
    Reviewed-by: Dawid Osuchowski <[email protected]>
    Signed-off-by: Karol Wachowski <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

accel/ivpu: Use threaded IRQ for IPC callback processing [+ + +]
Author: Karol Wachowski <[email protected]>
Date:   Tue Sep 29 08:07:24 2026 -0400

    accel/ivpu: Use threaded IRQ for IPC callback processing
    
    [ Upstream commit 85c9cc2d25f80534c1623621264018a655869cb2 ]
    
    Dispatching IPC callbacks from system_percpu_wq adds scheduling latency
    that is neither bounded nor predictable, which hurts job completion
    turnaround. Handle them from a threaded IRQ instead: the hard-IRQ
    handler drains the IPC FIFO and wakes the thread, which runs the
    callback consumers such as job-done processing.
    
    Job resource teardown can trigger IOMMU unmapping and context teardown,
    which is too slow to run from the IRQ thread. Defer it to a dedicated
    WQ_UNBOUND | WQ_MEM_RECLAIM workqueue via a per-device lockless list.
    UNBOUND keeps the long-running cleanup off the percpu workers and
    MEM_RECLAIM guarantees forward progress because the work frees buffer
    objects. The runtime PM reference taken at submission is released only
    after cleanup completes, otherwise runtime suspend could race the
    pending work and deadlock.
    
    Because cleanup is now asynchronous, userspace that rapidly recycles
    file descriptors or command queues can momentarily observe stale
    per-context resources and fail with -EMFILE or -EBUSY. Flush the
    cleanup work once and retry before giving up.
    
    Reviewed-by: Andrzej Kacprowski <[email protected]>
    Signed-off-by: Karol Wachowski <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Stable-dep-of: 72b782097e53 ("accel/ivpu: Use separate flag for job timeout")
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
af_packet: fix integer overflow in prb_calc_retire_blk_tmo() [+ + +]
Author: Dairui Zhang <[email protected]>
Date:   Wed Sep 23 13:01:01 2026 +0800

    af_packet: fix integer overflow in prb_calc_retire_blk_tmo()
    
    commit 56d82862a0a243ac14ba11b6d7b57ddc2d064b95 upstream.
    
    prb_calc_retire_blk_tmo() computes in 32-bit int arithmetic:
    
            mbits = (blk_size_in_bytes * 8) / (1024 * 1024);
    
    If I'm reading the validation right, tp_block_size is user
    controlled and packet_set_ring() only rejects values that are <= 0
    as int or not page aligned, so a 256MiB block goes right through
    (and alloc_one_pg_vec_page() even has a vzalloc fallback for it).
    0x10000000 * 8 wraps to INT_MIN, and on a NIC reporting 1 Gbps
    (div == 1) the function ends up returning -2047.
    
    The condition is actually (8 * size) mod 2^32 >= 2^31 && div == 1,
    so the trigger set is [256,512), [768,1024), [1280,1536) and
    [1792,2048) MiB. Other sizes wrap to non-negative values and faster
    links divide the unsigned value back below 2^31, which is why this
    doesn't blow up for everyone.
    
    What makes it fatal is what happens next in init_prb_bdqc():
    
            p1->interval_ktime = ms_to_ktime(prb_calc_retire_blk_tmo(...));
            hrtimer_start(&p1->retire_blk_timer, p1->interval_ktime,
                          HRTIMER_MODE_REL_SOFT);
    
    A negative relative timeout expires immediately. The callback
    unconditionally returns HRTIMER_RESTART, and hrtimer_forward() turns
    the negative interval into hrtimer_resolution:
    
            if (interval < hrtimer_resolution)
                    interval = hrtimer_resolution;
    
    So the SOFT timer re-fires at the maximum rate forever, holding
    sk_receive_queue.lock each pass. One CPU spins in softirq until the
    socket is closed. Repeat with more rings and the machine is gone.
    
    The overflow itself is ancient - it was introduced together with
    TPACKET_V3 in f6fb8f100b80 ("af-packet: TPACKET_V3 flexible buffer
    implementation."). Its effect prior to f7460d2989fa ("net:
    af_packet: Use hrtimer to do the retire operation", v6.18) was not
    as clear-cut, though: the return value was stored into an unsigned
    short retire_blk_tov, so a negative result was truncated, and a
    0-jiffy delay loop could be programmed as well. Neither is nearly
    as detrimental as the immediate maximum-rate spin the hrtimer
    conversion turned it into.
    
    (Unrelated to CVE-2019-20812 - that one was the ethtool failure path
    returning 0, which now returns DEFAULT_PRB_RETIRE_TOV.)
    
    Reproducer, needs CAP_NET_RAW (a --network host container has it by
    default) and a 1 Gbps NIC (QEMU e1000 works):
    
            int fd = socket(AF_PACKET, SOCK_RAW, htons(ETH_P_ALL));
            bind(fd, ...);
            int v = TPACKET_V3;
            setsockopt(fd, SOL_PACKET, PACKET_VERSION, &v, sizeof(v));
            struct tpacket_req3 req = {
                    .tp_block_size = 0x10000000,
                    .tp_block_nr = 1,
                    .tp_frame_size = 2048,
                    .tp_frame_nr = 0x10000000 / 2048,
                    .tp_retire_blk_tov = 0,
            };
            setsockopt(fd, SOL_PACKET, PACKET_RX_RING, &req, sizeof(req));
    
    Compute in 64 bits instead. The operands are already bounded by the
    existing validation, so nothing else changes. If you'd prefer a
    different fix, just say so and I'll respin.
    
    Fixes: f6fb8f100b80 ("af-packet: TPACKET_V3 flexible buffer implementation.")
    Cc: [email protected]
    Signed-off-by: Dairui Zhang <[email protected]>
    Reviewed-by: Willem de Bruijn <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
ALSA: hda/realtek: Add StarFighter HDA SSID [+ + +]
Author: Sean Rhodes <[email protected]>
Date:   Fri Jul 31 22:13:09 2026 +0100

    ALSA: hda/realtek: Add StarFighter HDA SSID
    
    [ Upstream commit cd401c70df472d3eddd0b6b055726a03c212181a ]
    
    Support the new StarFighter HDA SSID while keeping the existing SSID chained to the same quirk until the new match reaches backports.
    
    Signed-off-by: Sean Rhodes <[email protected]>
    Signed-off-by: Takashi Iwai <[email protected]>
    Link: https://patch.msgid.link/06865eaedf3de8dff199e9aa7e86cd135572f20f.1785532385.git.sean@starlabs.systems
    Signed-off-by: Sasha Levin <[email protected]>

ALSA: hda/realtek: Limit Star Labs internal mic boost [+ + +]
Author: Sean Rhodes <[email protected]>
Date:   Fri Jul 31 22:13:08 2026 +0100

    ALSA: hda/realtek: Limit Star Labs internal mic boost
    
    [ Upstream commit 186d4adbb40138e7cb7cffc87a81e95630ced123 ]
    
    The 30 dB internal mic boost is too high for laptops, especially with fans. Limit Star Labs internal mic boost to 10 dB.
    
    Signed-off-by: Sean Rhodes <[email protected]>
    Signed-off-by: Takashi Iwai <[email protected]>
    Link: https://patch.msgid.link/be87292613b24150d6321adac102b4b25d00e9e6.1785532385.git.sean@starlabs.systems
    Signed-off-by: Sasha Levin <[email protected]>

 
arm64/boot: Disable trapping of PMZR_EL0 writes to EL2 [+ + +]
Author: Fuad Tabba <[email protected]>
Date:   Tue Sep 22 19:14:30 2026 +0100

    arm64/boot: Disable trapping of PMZR_EL0 writes to EL2
    
    commit 2bc6b218717b9d08f466f88209251d54bc09b207 upstream.
    
    __init_el2_fgt2() writes one mask to both HDFGRTR2_EL2 and HDFGWTR2_EL2.
    PMZR_EL0 is write-only, so its trap bit, nPMZR_EL0, exists only in
    HDFGWTR2_EL2 and is therefore never set: a PMZR_EL0 write from the host
    traps to EL2, where the nVHE hypervisor has no handler and BUG()s. The
    kernel never writes PMZR_EL0, but kernel.perf_user_access=1 has the PMU
    driver set PMUSERENR_EL0.UEN for a task with a user-read event, so a
    write from EL0 reaches the trap and takes the host down without a panic
    message.
    
    Accumulate the HDFGWTR2_EL2 bits separately, as __init_el2_fgt() already
    does for HDFGWTR_EL2, and set nPMZR_EL0 with the other FEAT_PMUv3p9
    bits.
    
    Fixes: 858c7bfcb35e1 ("arm64/boot: Enable EL2 requirements for FEAT_PMUv3p9")
    Cc: [email protected]
    Signed-off-by: Fuad Tabba <[email protected]>
    Reviewed-by: Anshuman Khandual <[email protected]>
    Reviewed-by: Oliver Upton <[email protected]>
    Signed-off-by: Will Deacon <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
arm64: errata: match the target implementation CPU's own MIDR [+ + +]
Author: David Carlier <[email protected]>
Date:   Sun Sep 6 13:14:16 2026 +0100

    arm64: errata: match the target implementation CPU's own MIDR
    
    commit b7403afb7a5f85073243df10238b3483958ad69e upstream.
    
    __is_affected_midr_range() is handed the MIDR and REVIDR of one target
    implementation CPU, but tests the erratum's range with is_midr_in_range(),
    which re-scans all of target_impl_cpus[] and ignores the @midr argument.
    The range test is thus constant across the per-CPU loop in
    is_affected_midr_range() and only answers "is any target CPU in range".
    
    Since just the fixed_revs REVIDR check uses the iteration's own registers,
    an out-of-range target CPU can decide whether a MIDR_FIXED() exemption
    applies. A VM then enables a workaround whose only in-range CPU is fixed
    silicon, e.g. erratum 2658417 on a Cortex-A510 r1p1 with REVIDR_EL1[25]
    set.
    
    Factor the range test into __is_midr_in_range(), which takes an explicit
    MIDR, and use it in __is_affected_midr_range().
    
    Fixes: 86edf6bdcf05 ("smccc/kvm_guest: Enable errata based on implementation CPUs")
    Cc: [email protected]
    Assisted-by: Claude:claude-opus-5
    Signed-off-by: David Carlier <[email protected]>
    Signed-off-by: Will Deacon <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

arm64: io: Reject non-user protection in ioremap_prot() [+ + +]
Author: Zeng Heng <[email protected]>
Date:   Fri Sep 11 09:58:59 2026 +0800

    arm64: io: Reject non-user protection in ioremap_prot()
    
    [ Upstream commit bb756b11ad63832ebee58caf9e8f9381eaecff9f ]
    
    Mapping a stack-top page via /dev/mem with PROT_NONE and then
    reading that process's /proc/<pid>/cmdline triggers a spurious WARN
    in ioremap_prot() through generic_access_phys():
    
      WARNING: ./arch/arm64/include/asm/io.h:275 at generic_access_phys
      Call trace:
        generic_access_phys+0x1c8/0x228 (P)
        __access_remote_vm+0x2b4/0x398
        access_remote_vm+0x14/0x30
        get_mm_cmdline+0xf8/0x2a0
        proc_pid_cmdline_read+0x68/0x120
    
    generic_access_phys() passes the protection derived from the user PTE
    to ioremap_prot(). On arm64, a PROT_NONE mapping is represented by a
    present-invalid PTE, so pte_present() still returns true and the
    protection reaches ioremap_prot().
    
    A PROT_NONE mapping does not have PTE_USER, causing the existing
    WARN_ON_ONCE() in ioremap_prot() to fire even though this is a valid
    user mapping. Execute-only mappings have the same issue and must not
    be readable through this path either.
    
    ioremap_prot() should therefore reject protection values without
    PTE_USER without warning. This makes the access fail cleanly for
    PROT_NONE and execute-only mappings while retaining the existing
    user-protection contract.
    
    Fixes: 8f098037139b ("arm64: io: Extract user memory type in ioremap_prot()")
    Signed-off-by: Zeng Heng <[email protected]>
    Reviewed-by: Catalin Marinas <[email protected]>
    Signed-off-by: Will Deacon <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
ata: libata-scsi: bound the ATA passthru sense descriptor writes [+ + +]
Author: Matthias Goergens <[email protected]>
Date:   Thu Sep 24 01:52:03 2026 +0800

    ata: libata-scsi: bound the ATA passthru sense descriptor writes
    
    commit 80320b278fea07ffcda3f57b67b61658e0a4e1ca upstream.
    
    When an ATA PASS-THROUGH command to an ATAPI device fails, the sense
    buffer holds the device's REQUEST SENSE reply, and
    ata_scsi_set_passthru_sense_fields() trusts its additional length
    byte, sb[7], when adding the ATA Status Return descriptor.  A faulty
    or malicious device can use that to make the kernel read and write
    past the 96-byte buffer in three ways:
    
    - scsi_sense_desc_find() is passed sb[7] + 8 as the buffer length, so
      its clamp against sb[7] does nothing and the walk runs off the end.
    - A type-9 descriptor found near the end is filled in unchecked.
    - A new descriptor at sb[8 + len] needs len + 22 bytes, not len + 14,
      so len 75..82 writes up to 8 bytes past the end.
    
    Reproduced with KASAN under qemu, with the emulated ATAPI REQUEST SENSE
    reply patched:
    
      BUG: KASAN: slab-out-of-bounds in scsi_sense_desc_find+0x1a5/0x210
      BUG: KASAN: slab-out-of-bounds in ata_scsi_qc_complete+0x1a15/0x1a50
    
    Both are gone with this patch, and a valid descriptor is still filled
    in.
    
    Fixes: 97981926224a ("ata: libata-scsi: Do not overwrite valid sense data when CK_COND=1")
    Cc: [email protected]
    Reviewed-by: Damien Le Moal <[email protected]>
    Signed-off-by: Matthias Goergens <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Niklas Cassel <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
autofs: fix sbi->pipe file reference leak in autofs_kill_sb() [+ + +]
Author: Hui Peng <[email protected]>
Date:   Sat Sep 19 20:48:08 2026 +0000

    autofs: fix sbi->pipe file reference leak in autofs_kill_sb()
    
    [ Upstream commit aa5e44b29ffe4eaa08cc2237fd65bc2596bc023e ]
    
    When autofs_fill_super() fails before clearing AUTOFS_SBI_CATATONIC (for
    example, when find_get_pid() fails on an invalid pgrp mount option, or
    when an fs_context is closed before mounting), deactivate_locked_super()
    invokes autofs_kill_sb() -> autofs_catatonic_mode(sbi).
    
    Because AUTOFS_SBI_CATATONIC is still set in sbi->flags,
    autofs_catatonic_mode() returns early without calling fput(sbi->pipe),
    permanently leaking the pipe struct file reference.
    
    Explicitly release sbi->pipe in autofs_kill_sb() if it is still non-NULL
    after autofs_catatonic_mode().
    
    Fixes: ebc921ca9b92 ("autofs: copy autofs4 to autofs")
    Signed-off-by: Hui Peng <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Christian Brauner (Amutable) <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
Bluetooth: bnep: fix out-of-bounds reads on short RX/TX frames and control fallthrough [+ + +]
Author: Hui Peng <[email protected]>
Date:   Sat Sep 19 22:17:38 2026 +0000

    Bluetooth: bnep: fix out-of-bounds reads on short RX/TX frames and control fallthrough
    
    [ Upstream commit f0ca020cbb9bb7f3f4ea8ba1dfcf30a282aec91e ]
    
    Fix multiple out-of-bounds reads in Bluetooth BNEP frame processing:
    
    1. In bnep_rx_frame() and bnep_ctrl_frame() (net/bluetooth/bnep/core.c),
       use pskb_may_pull() to verify the BNEP header, control type byte,
       filter count, and extension headers exist before reading them, and
       return 0 after handling BNEP_CONTROL instead of falling through to
       Ethernet frame submission when no extension headers follow.
    2. In bnep_net_xmit() (net/bluetooth/bnep/netdev.c), verify skb->len >=
       ETH_HLEN with pskb_may_pull() before reading the 14-byte Ethernet
       header to prevent an out-of-bounds heap read and infoleak on short
       AF_PACKET TX frames.
    
    Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
    Assisted-by: LLM
    Signed-off-by: Hui Peng <[email protected]>
    Signed-off-by: Luiz Augusto von Dentz <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

Bluetooth: btintel_pcie: validate device-supplied DMA indices [+ + +]
Author: Ravindra <[email protected]>
Date:   Tue Sep 15 10:42:15 2026 +0530

    Bluetooth: btintel_pcie: validate device-supplied DMA indices
    
    [ Upstream commit 37a11129345337efd6eef8e62b03b6348cd0dd8b ]
    
    In btintel_pcie_msix_rx_handle(), the driver processes RX completion
    descriptors (urbd1) written by the PCIe device into DMA-coherent memory.
    urbd1->frbd_tag (a 16-bit field fully controlled by the device firmware
    via DMA) is used directly as an array index into rxq->bufs[] without any
    bounds check. rxq->bufs[] has only BTINTEL_PCIE_RX_DESCS_COUNT (64)
    entries, while frbd_tag can be any value 0-65535. A malicious or
    malfunctioning device can write an out-of-range frbd_tag, causing the
    driver to dereference an out-of-bounds data_buf pointer.
    
    Additionally, cr_hia is read from a DMA-shared index array also writable
    by the device; if the device sets cr_hia >= rxq->count, the while-loop
    never terminates because cr_tia is wrapped via modulo rxq->count and can
    never equal an out-of-range cr_hia.
    
    Add bounds validation for cr_hia and frbd_tag in the RX path, and cr_hia
    in the TX path. Log invalid values with bt_dev_err before returning.
    
    Fixes: c2b636b3f788 ("Bluetooth: btintel_pcie: Add support for PCIe transport")
    Signed-off-by: Ravindra <[email protected]>
    Signed-off-by: Luiz Augusto von Dentz <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

Bluetooth: btnxpuart: Fix skb leak in nxp_process_fw_dump() [+ + +]
Author: Zijun Hu <[email protected]>
Date:   Tue Sep 15 19:17:18 2026 -0700

    Bluetooth: btnxpuart: Fix skb leak in nxp_process_fw_dump()
    
    [ Upstream commit f2bbb36426581045a8bf7793da5419b9375e4348 ]
    
    When CONFIG_DEV_COREDUMP=n, hci_devcd_append() returns -EOPNOTSUPP
    without freeing its skb argument. This leaks the cloned skb and also
    prevents nxp_set_ind_reset() from being called to perform recovery.
    
    Fix by guarding the hci_devcd_append(hdev, skb_clone(skb, GFP_ATOMIC))
    call with IS_ENABLED(CONFIG_DEV_COREDUMP).
    
    Fixes: 998e447f443f ("Bluetooth: btnxpuart: Add support for HCI coredump feature")
    Signed-off-by: Zijun Hu <[email protected]>
    Signed-off-by: Luiz Augusto von Dentz <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

Bluetooth: hci_conn: fix CIS hold ownership on reuse [+ + +]
Author: Aldo Ariel Panzardo <[email protected]>
Date:   Tue Sep 15 13:04:29 2026 -0300

    Bluetooth: hci_conn: fix CIS hold ownership on reuse
    
    commit e06d549fcd4a0ba381ed67ddf1ab3c7a6ca4314c upstream.
    
    Commit 69997d50ec57 ("Bluetooth: ISO: handle bound CIS cleanup via
    hci_conn") made hci_bind_cis() and hci_connect_cis() return a
    connection with one hold for the ISO layer.  hci_bind_cis() currently
    takes that hold only after configuring a CIS, so its BT_CONNECTED and
    matching BT_BOUND paths return a bare lookup result.  Its configuration
    failure path can likewise call hci_conn_drop() before taking a hold.
    
    Take the hold before any state-dependent return or configuration error
    so every successful return follows the documented ownership contract
    and every error drop is balanced.
    
    hci_connect_cis() also assumes hci_conn_link() always takes a new CIS
    hold before dropping the one returned by hci_bind_cis().  However, the
    helper returns an existing link without taking another hold.  In that
    case, preserve the CIS hold for the caller and drop the redundant LE
    hold because the existing link already owns its parent hold.  Returning
    early also avoids changing an existing CIS back to BT_CONNECT.
    
    Fixes: 69997d50ec57 ("Bluetooth: ISO: handle bound CIS cleanup via hci_conn")
    Cc: [email protected]
    Signed-off-by: Aldo Ariel Panzardo <[email protected]>
    Signed-off-by: Luiz Augusto von Dentz <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

Bluetooth: hci_sock: reject out-of-range OCF values [+ + +]
Author: Aldo Ariel Panzardo <[email protected]>
Date:   Tue Sep 15 13:03:58 2026 -0300

    Bluetooth: hci_sock: reject out-of-range OCF values
    
    commit e93fad891c72deb84cae49430163b384ebcc92b1 upstream.
    
    The raw HCI socket security filter has 128 OCF bits per supported OGF,
    but masks the 10-bit OCF with 127 before looking up the command. An
    unprivileged socket can therefore submit a reserved OCF that aliases an
    allowlisted command modulo 128.
    
    A conforming controller should reject reserved opcodes. Nevertheless,
    the security decision must apply to the opcode that will actually be
    sent, especially since controller-specific behavior is outside the host
    stack's control.
    
    Reject OCF values that cannot be represented by the security filter
    instead of aliasing them onto an unrelated command.
    
    Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
    Cc: [email protected]
    Signed-off-by: Aldo Ariel Panzardo <[email protected]>
    Signed-off-by: Luiz Augusto von Dentz <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

Bluetooth: hci_sock: validate event length before filtering [+ + +]
Author: Aldo Ariel Panzardo <[email protected]>
Date:   Tue Sep 15 13:03:07 2026 -0300

    Bluetooth: hci_sock: validate event length before filtering
    
    commit b0a6cf99afd57a39598b1beca0e86ef5004980de upstream.
    
    is_filtered_packet() reads the event code from skb->data[0] without first
    checking that the skb is nonempty. When an opcode filter is configured,
    it also reads the command opcode at offsets 3 or 4 without checking that
    a Command Complete or Command Status event is long enough.
    
    hci_send_to_sock() invokes the filter before hci_event_packet() validates
    the event header. A malformed event supplied by a controller or a vhci
    device can therefore cause an out-of-bounds read.
    
    Keep the unmasked event code for the opcode checks. The masked value is
    needed for the 64-bit event bitmap, but using it to identify command events
    aliases event codes above 0x3f. In particular, Synchronous Train Complete
    (0x4f) was treated as Command Status (0x0f) even though its payload has no
    opcode.
    
    Reject actual command events that are too short for the field being
    inspected. A truncated command event cannot match a configured opcode.
    
    Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
    Cc: [email protected]
    Signed-off-by: Aldo Ariel Panzardo <[email protected]>
    Signed-off-by: Luiz Augusto von Dentz <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

Bluetooth: ISO: balance the parent hold in hci_bind_bis() [+ + +]
Author: Aldo Ariel Panzardo <[email protected]>
Date:   Tue Sep 15 13:03:32 2026 -0300

    Bluetooth: ISO: balance the parent hold in hci_bind_bis()
    
    commit 4c94557dd02569efa6c1072a0439addaef9a5224 upstream.
    
    hci_conn_link() takes a lifetime reference to its parent with
    hci_conn_get(), but only takes an operational hold on the child.
    hci_conn_unlink() later balances both a hold and a reference on the
    parent.
    
    The SCO and CIS paths pass a parent acquired from a connect helper, so
    it already has a hold. For an additional BIS, hci_bind_bis() obtains the
    parent from hci_conn_hash_lookup_big(), which returns a bare pointer.
    Unlinking the child then drops the parent's existing hold and can
    schedule it for disconnection while its socket is still using it.
    
    Take a hold on the parent before linking it and drop that hold if linking
    fails. A successful link transfers the hold to hci_conn_unlink().
    
    Fixes: fa224d0c094a ("Bluetooth: ISO: Reassociate a socket with an active BIS")
    Cc: [email protected]
    Signed-off-by: Aldo Ariel Panzardo <[email protected]>
    Signed-off-by: Luiz Augusto von Dentz <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

Bluetooth: ISO: release unused CIS holds after channel attach [+ + +]
Author: Aldo Ariel Panzardo <[email protected]>
Date:   Tue Sep 15 13:04:30 2026 -0300

    Bluetooth: ISO: release unused CIS holds after channel attach
    
    commit 0fcd4dad555c96e0bd3a1b8c569f989be85c7341 upstream.
    
    hci_bind_cis() and hci_connect_cis() return one hci_conn hold for the
    ISO layer.  A new channel association consumes that hold, which is
    eventually released by iso_conn_free().
    
    There are two cases where iso_chan_add() does not create an association:
    it returns success when the socket is already attached to the same
    iso_conn, and it returns -EBUSY when another socket is attached.  The
    hold returned for the current call is unused in both cases.  This occurs
    when deferred setup calls iso_connect_cis() again for its existing
    socket, or when another socket attempts to reuse the CIS.
    
    Detect the idempotent case while the connection is locked and release
    the unused hold after iso_chan_add().  Also release it on -EBUSY.  Do not
    drop it for other errors: a newly allocated iso_conn releases the
    transferred hold when its last temporary reference is put.
    
    Fixes: 69997d50ec57 ("Bluetooth: ISO: handle bound CIS cleanup via hci_conn")
    Cc: [email protected]
    Signed-off-by: Aldo Ariel Panzardo <[email protected]>
    Signed-off-by: Luiz Augusto von Dentz <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

Bluetooth: L2CAP: validate frame length before control and FCS access [+ + +]
Author: Aldo Ariel Panzardo <[email protected]>
Date:   Tue Sep 15 13:02:39 2026 -0300

    Bluetooth: L2CAP: validate frame length before control and FCS access
    
    commit 6c78a213d9070b610c7f418af2c25b66180b7e37 upstream.
    
    l2cap_data_rcv() unpacks either a two-byte or four-byte control field
    without first ensuring that it is present. A short ERTM or streaming-mode
    frame can therefore cause an out-of-bounds read.
    
    There is a second short-frame case when CRC16 is enabled. After the
    control field is pulled, l2cap_check_fcs() subtracts two from skb->len
    without checking it. If fewer than two bytes remain, the subtraction
    wraps; skb_trim() leaves the buffer unchanged and the subsequent FCS
    load reads past the logical end of the frame.
    
    Validate that the frame contains both its control field and, when
    enabled, its FCS before either field is accessed.
    
    Fixes: 1c2acffb76d4 ("Bluetooth: Add initial support for ERTM packets transfers")
    Fixes: fcc203c30d72 ("Bluetooth: Add support for FCS option to L2CAP")
    Cc: [email protected]
    Signed-off-by: Aldo Ariel Panzardo <[email protected]>
    Signed-off-by: Luiz Augusto von Dentz <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

Bluetooth: mgmt: Dequeue pending mesh_send_sync entries on cancel [+ + +]
Author: Lee Jones <[email protected]>
Date:   Tue Sep 15 12:08:22 2026 +0000

    Bluetooth: mgmt: Dequeue pending mesh_send_sync entries on cancel
    
    [ Upstream commit 71af682ba4692c2ed9ace4c3d4ca462ae368c029 ]
    
    In send_cancel(), pending mesh_tx objects are removed from the
    hdev->mesh_pending list and freed via mesh_send_complete().  However, if
    a mesh transmission was already queued onto hdev->cmd_sync_work_list via
    mesh_next(), the queued entry retains a raw pointer to mesh_tx.
    
    When hci_cmd_sync_work later processes the entry, it attempts to execute
    mesh_send_sync and its destroy callback mesh_send_start_complete using
    the already freed mesh_tx pointer, leading to a use-after-free.
    
    Fix this by invoking hci_cmd_sync_dequeue() for mesh_send_sync on the
    target mesh_tx before completing it.  If the entry is found and dequeued,
    its destroy callback will complete and free the object; otherwise,
    mesh_send_complete() is called directly.
    
    Additionally, ensure the transmission queue advances after cancellation
    or errors.  In mesh_send_start_complete(), call mesh_next() on error
    unless err is -ECANCELED, because hci_cmd_sync_dequeue() holds
    hdev->cmd_sync_work_lock and calling mesh_next() synchronously would
    deadlock.  Instead, advance the queue in send_cancel() once the lock is
    released and if no transmission is in progress.
    
    Fixes: b338d91703fa ("Bluetooth: Implement support for Mesh")
    Signed-off-by: Lee Jones <[email protected]>
    Signed-off-by: Luiz Augusto von Dentz <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

Bluetooth: mgmt: fix race in read_unconf_index_list() [+ + +]
Author: Aldo Ariel Panzardo <[email protected]>
Date:   Tue Sep 15 12:59:52 2026 -0300

    Bluetooth: mgmt: fix race in read_unconf_index_list()
    
    commit b5dbb41b212c50c095a4dbee3017a84fe94f033b upstream.
    
    read_unconf_index_list() counts unconfigured controllers before allocating
    its response, then checks the device flags again while filling it.
    
    hci_dev_list_lock stabilizes list membership, but it does not serialize the
    per-device flags. During asynchronous controller setup, the worker can set
    HCI_UNCONFIGURED and clear HCI_SETUP between the two passes. A controller
    omitted from the allocation count can then become eligible for the fill
    pass, causing an out-of-bounds write to rp->index[].
    
    Allocate space for every device on hci_dev_list. Since list membership
    cannot change while hci_dev_list_lock is held, the response remains large
    enough regardless of flag transitions. The reported count and response
    length still include only eligible unconfigured controllers.
    
    Fixes: 73d1df2a7a10 ("Bluetooth: Add support for Read Unconfigured Index List command")
    Cc: [email protected]
    Signed-off-by: Aldo Ariel Panzardo <[email protected]>
    Signed-off-by: Luiz Augusto von Dentz <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

Bluetooth: RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO [+ + +]
Author: Hui Peng <[email protected]>
Date:   Sat Sep 19 11:25:18 2026 +0000

    Bluetooth: RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO
    
    commit 46f8ffd0a1f1eb6cbc94946a92c11ef601e228a1 upstream.
    
    The RFCOMM_CONNINFO getsockopt handler accepts a socket that is not
    connected as long as deferred setup is enabled:
    
            if (sk->sk_state != BT_CONNECTED &&
                                    !rfcomm_pi(sk)->dlc->defer_setup) {
                    err = -ENOTCONN;
                    break;
            }
            l2cap_sk = rfcomm_pi(sk)->dlc->session->sock->sk;
    
    dlc->defer_setup is set in rfcomm_sock_init() when rfcomm_connect_ind()
    creates a child socket for an incoming connection on a listening socket
    that has BT_DEFER_SETUP enabled. It is never cleared afterwards. The
    session, however, can go away underneath it.
    
    rfcomm_recv_disc() forces the dlc state before tearing it down:
    
            d->state = BT_CLOSED;
            __rfcomm_dlc_close(d, err);
    
    The RFCOMM_DEFER_SETUP early return in __rfcomm_dlc_close() only covers
    BT_CONNECT, BT_CONFIG, BT_OPEN and BT_CONNECT2, so with the state
    already BT_CLOSED that switch does not match and the function falls
    through to rfcomm_dlc_unlink(), which sets d->session = NULL, while
    d->defer_setup stays 1.
    
    A getsockopt(SOL_RFCOMM, RFCOMM_CONNINFO) on the accepted socket after
    that point therefore skips the -ENOTCONN path -- sk->sk_state is
    BT_CLOSED, but dlc->defer_setup is still set -- and dereferences the
    NULL session. No race is needed: once the DISC has been processed, the
    dereference is unconditional.
    
    Reproduced on a KASAN kernel under QEMU with a BR/EDR peer emulated over
    /dev/vhci: the peer brings up an ACL link, opens L2CAP on the RFCOMM
    PSM, starts a session and sends SABM for a channel bound with
    BT_DEFER_SETUP, and sends DISC for that dlci after the socket has been
    accepted. getsockopt(SOL_RFCOMM, RFCOMM_CONNINFO) on the accepted
    socket then hits:
    
     Oops: general protection fault, probably for non-canonical address
     0xdffffc0000000002: 0000 [#1] SMP KASAN PTI
     KASAN: null-ptr-deref in range [0x0000000000000010-0x0000000000000017]
     CPU: 1 UID: 0 PID: 150 Comm: init Tainted: G B 7.3.0-rc3-g5dd1818b15d9
     Hardware name: QEMU Standard PC (i440FX + PIIX, 1996)
     RIP: 0010:rfcomm_sock_getsockopt+0x529/0x780
     Call Trace:
      <TASK>
      do_sock_getsockopt+0x3ad/0x7d0
      __sys_getsockopt+0x10e/0x1b0
      __x64_sys_getsockopt+0xc2/0x160
      do_syscall_64+0xda/0x4b0
      entry_SYSCALL_64_after_hwframe+0x77/0x7f
      </TASK>
    
    0x10 is the offset of sock in struct rfcomm_session;
    rfcomm_sock_getsockopt_old() is inlined into rfcomm_sock_getsockopt().
    
    Commit 43a556b2fd43 ("Bluetooth: RFCOMM: take rfcomm_mutex for the
    deferred setup accept") fixed the same "a remote DISC clears the session
    while deferred setup is still flagged" problem in rfcomm_dlc_accept();
    this is the remaining instance of it, in the getsockopt path.
    
    Deferred setup only leaves a socket usable here once it has reached
    BT_CONNECT2, so restrict the exception to that state and check that a
    session is actually present before following it.
    
    Fixes: bb23c0ab8246 ("Bluetooth: Add support for deferring RFCOMM connection setup")
    Cc: [email protected]
    Assisted-by: LLM
    Signed-off-by: Hui Peng <[email protected]>
    Signed-off-by: Luiz Augusto von Dentz <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

Bluetooth: RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame() [+ + +]
Author: Hui Peng <[email protected]>
Date:   Sat Sep 19 11:25:14 2026 +0000

    Bluetooth: RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame()
    
    [ Upstream commit 6d91041bb38b97e2feb625123cc0529d7b83a0e1 ]
    
    While rfcomm_recv_frame() verifies that skb->len is at least
    sizeof(*hdr) + 1 (4 bytes: 3-byte header + 1-byte FCS), an RFCOMM frame
    with an extended 2-byte length field (!__test_ea(hdr->len)) has a 4-byte
    header plus a 1-byte FCS (5 bytes minimum, sizeof(*hdr) + 2).
    
    When a 4-byte RFCOMM frame with EA == 0 arrives:
    1. The initial skb->len < sizeof(*hdr) + 1 check passes (4 < 4 is false).
    2. Trimming the FCS byte decrements skb->len to 3.
    3. If __check_fcs() succeeds, skb_pull(skb, 4) fails (4 > 3) and returns
       NULL without advancing skb->data.
    4. Because the return value of skb_pull() is ignored, the un-pulled
       3-byte struct rfcomm_hdr remains at skb->data and is either queued as
       application payload via rfcomm_recv_data() or parsed as a multiplexer
       control command via rfcomm_recv_mcc() on DLCI 0.
    
    Fix this by extending the length check in rfcomm_recv_frame() to also
    require skb->len >= sizeof(*hdr) + 2 when !__test_ea(hdr->len).
    
    Fixes: b230e5bf501c ("Bluetooth: RFCOMM: validate skb length in rfcomm_recv_frame")
    Assisted-by: LLM
    Signed-off-by: Hui Peng <[email protected]>
    Signed-off-by: Luiz Augusto von Dentz <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

Bluetooth: SMP: reject Security Request over BR/EDR [+ + +]
Author: Christiano Amora <[email protected]>
Date:   Wed Sep 16 10:46:22 2026 -0300

    Bluetooth: SMP: reject Security Request over BR/EDR
    
    [ Upstream commit f033482d76a9f18080c7a40c5f9c678bd7adc8f3 ]
    
    Bose QC Ultra Headphones (dual-mode, same public address on both
    transports) occasionally send an SMP Security Request on the BR/EDR
    SMP fixed channel right after the ACL link is encrypted. The kernel
    handles it as if it were an LE link: smp_cmd_security_req() has no
    transport check, smp_ltk_encrypt() looks up an LTK with the ACL
    connection's dst_type, and hci_find_ltk() matches the peer's LE LTK
    because the LE public address type is stored as ADDR_LE_DEV_PUBLIC (0),
    the same value as BDADDR_BREDR. HCI_OP_LE_START_ENC is then issued on
    the ACL handle, the controller rejects it with Invalid HCI Command
    Parameters, and hci_cs_le_start_enc() disconnects the link with
    HCI_ERROR_AUTH_FAILURE. The headphones drop within a second of
    connecting, before any profile is up; a manual reconnect works.
    
    btmon (MediaTek MT7922, kernel 7.0.12):
    
      > HCI Event: Encryption Change (0x08) plen 4
            Status: Success (0x00)
            Handle: 50 Address: BC:87:FA:47:73:5E (Bose Corporation)
            Encryption: Enabled with AES-CCM (0x02)
      > ACL Data RX: Handle 50 flags 0x02 dlen 6
            BR/EDR SMP: Security Request (0x0b) len 1
            Authentication requirement: No bonding, No MITM, SC (0x08)
      < HCI Command: LE Start Encryption (0x08|0x0019) plen 28
            Handle: 50 Address: BC:87:FA:47:73:5E (Bose Corporation)
      > HCI Event: Command Status (0x0f) plen 4
            LE Start Encryption (0x08|0x0019) ncmd 1
            Status: Invalid HCI Command Parameters (0x12)
      < HCI Command: Disconnect (0x01|0x0006) plen 3
            Handle: 50 Address: BC:87:FA:47:73:5E (Bose Corporation)
            Reason: Authentication Failure (0x05)
    
    SMP over BR/EDR is limited to cross-transport key derivation; the
    Security Request procedure (Core Specification Vol 3, Part H, Section
    2.4.6, PDU in Section 3.6.7) has no BR/EDR counterpart. Reply with
    Pairing Failed / Command Not Supported on a non-LE link, before the PDU
    is parsed, and keep the connection. The reply is sent directly rather
    than through smp_failure(): rejecting a command on the wrong transport
    is not an authentication failure, and MGMT_EV_AUTH_FAILED would make
    bluetoothd disconnect the device.
    
    Tested on the affected host (kernel 7.0.12, MediaTek MT7922, Bose QC
    Ultra) with the patched module built out of tree: 7 days and 49
    reconnects without a drop, against 2 drops in the 3 days before the
    patch. Every disconnect in that week had a userspace or remote reason.
    
    Fixes: b5ae344d4c0f ("Bluetooth: Add full SMP BR/EDR support")
    Assisted-by: LLM
    Signed-off-by: Christiano Amora <[email protected]>
    Signed-off-by: Luiz Augusto von Dentz <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
bna: prevent IOC timer rearm during teardown [+ + +]
Author: Myeonghun Pak <[email protected]>
Date:   Mon Sep 21 21:46:05 2026 -0400

    bna: prevent IOC timer rearm during teardown
    
    commit 77b1718e39e5c9f6956fb60807326af20baf889d upstream.
    
    bna: prevent IOC timer rearm during teardown
    
    bnad_pci_remove() and the probe disable_ioceth path call
    timer_delete_sync() for ioc_timer, sem_timer and hb_timer, but not for
    iocpf_timer.  bnad_iocpf_timeout() then takes bnad->bna_lock after
    free_netdev() has freed the struct bnad.
    
    Deleting iocpf_timer last does not fix this.  sem_timer and
    iocpf_timer rearm each other: bnad_iocpf_sem_timeout() can arm
    iocpf_timer, and bnad_iocpf_timeout() arms sem_timer from
    bfa_ioc_hw_sem_get() when the semaphore is busy.
    timer_delete_sync() only waits out its own callback.
    bnad_ioceth_disable() can time out and leave that callback live.
    
    Shut all four IOC timers down with timer_shutdown_sync() on both
    paths, so a later mod_timer() is ignored.
    
    This issue was identified during our ongoing static-analysis research
    while reviewing kernel code.
    
    Fixes: 1d32f7696286 ("bna: IOC failure auto recovery fix")
    Cc: [email protected]
    Assisted-by: LLM
    Co-developed-by: Ijae Kim <[email protected]>
    Signed-off-by: Ijae Kim <[email protected]>
    Signed-off-by: Myeonghun Pak <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
bonding: crypto offload enabled, non-offload slave failover, rekey failed [+ + +]
Author: David Dai <[email protected]>
Date:   Fri Sep 18 16:11:55 2026 -0500

    bonding: crypto offload enabled, non-offload slave failover, rekey failed
    
    [ Upstream commit 00efbbd40bd5fd92c67b7cf1aab8904fa59a96f6 ]
    
    Create a bonding device (i.e. bond0) in active-backup mode, 2 slaves.
    Active slave: offload capable interface (i.e. eth1), primary interface.
    Backup slave: non-offload capable interface(i.e. eth2).
    Configure strongswan service swantl.conf child SA "hw_offload = crypto"
    Start strongswan service
    IPSec Crytpo Offload is enabled on top of bond0. i.e.
    ip xfrm state |grep offload
            crypto offload parameters: dev bond0 dir out mode crypto
            crypto offload parameters: dev bond0 dir in mode crypto
    
    Active slave eth1 takes adavantage of IPSec Crypto Offload capability.
    
    If active slave eth1 is down for any reason (i.e. eth1 link down):
    ip link set down dev eth1
    non-offload capable interface eth2 failover to becomes active slave.
    The existing SAs can continue use software IPsec after failover.
    Traffic still keeps going properly.
    
    However if eth1 link had not recovered yet, strongswan service does
    new child SA rekey, or uses swanctl command to do new child SA rekey,
    it will fail because active slave eth2 doesn't support crypto offload.
    In bond_ipsec_add_sa routine, it returns -EINVAL now, which is
    treated as fatal error by xfrm_dev_state_add routine in kernel xfrm.
    
    To make the non-offload active slave survive the child SA rekey, need
    to make bond_ipsec_add_sa routine returns -EOPNOTSUPP instead when
    active slave doesn't support IPsec Crypto offload, the xfrm will
    gracefully fallback to create new SA using Software IPsec.
    Network traffic can keep going.
    
    After offload capable interface eth1 link is up, becomes active slave,
    next time strongswan child SA rekey will create a new SA which enables
    crypto offload again.
    
    Fixes: 18cb261afd7b ("bonding: support hardware encryption offload to slaves")
    Signed-off-by: David Dai <[email protected]>
    Reviewed-by: Hangbin Liu <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
bpf, arm64: set up the frame pointer for the exception callback [+ + +]
Author: Donggeun Yoo <[email protected]>
Date:   Mon Sep 7 22:06:23 2026 +0900

    bpf, arm64: set up the frame pointer for the exception callback
    
    [ Upstream commit ef1fb82f12186dd26153b14d9fbcf4ec98db81b3 ]
    
    A program acting as exception boundary saves all callee-saved registers,
    so build_prologue() takes the exception_cb path and never calls
    push_callee_regs(). That is the only place find_used_callee_regs() runs,
    and with it the only place ctx->fp_used is set, so the callback prologue
    does not emit the
    
      mov x25, sp
    
    that points BPF_REG_FP at the frame the callback runs on. x25 keeps
    whatever it held when bpf_throw() was called. If the throw came from a
    subprogram that uses its own BPF stack, that is the subprogram's frame
    pointer, and since the subprogram never returns it never restores x25
    either.
    
    Stack accesses through BPF_REG_FP are rewritten to be stack pointer
    relative, so those still land in the callback's own frame. Materializing
    the register does not: a callback that passes the address of a local
    variable to a helper hands over an address in the dead subprogram's
    frame. That address is below the callback's stack pointer by then, and
    the helper's own call chain covers it, so the helper can write over its
    own return address. 0x1234 below is the value the helper was asked to
    store:
    
      pc : 0x1234
      lr : 0x1234
      Call trace:
       0x1234 (P)
       bpf_test_run+0x188/0x3e0
       bpf_prog_test_run_skb+0x47c/0x998
       __sys_bpf+0xbdc/0xdd8
      Kernel panic - not syncing: Oops: Fatal exception in interrupt
    
    Set ctx->fp_used on the exception callback path so that the existing code
    further down sets x25 from the stack pointer. The epilogue restores it
    from the main program's save area along with the other callee-saved
    registers, as it already does. x86 sets the frame pointer for the
    callback from the argument it is passed, and powerpc computes it from
    the stack pointer.
    
    Fixes: 5d4fa9ec5643 ("bpf, arm64: Avoid blindly saving/restoring all callee-saved registers")
    Acked-by: Xu Kuohai <[email protected]>
    Signed-off-by: Donggeun Yoo <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
bpf, sockmap: Fix self-redirect copied_seq double-counting [+ + +]
Author: Geliang Tang <[email protected]>
Date:   Tue Sep 8 17:08:32 2026 +0800

    bpf, sockmap: Fix self-redirect copied_seq double-counting
    
    [ Upstream commit 490a83d6386eec1d29f470c8d7331677fb46c3b7 ]
    
    When a BPF stream_verdict program redirects an skb back to the same
    socket (self-redirect with BPF_F_INGRESS), sk_psock_verdict_apply()
    calls tcp_eat_skb() which advances tcp_sk->copied_seq. However, the
    skb is then delivered to the socket's psock ingress queue and later
    read by tcp_bpf_recvmsg_parser(), which also advances copied_seq via
    the copied_from_self accounting path. This double-counting causes
    copied_seq to advance by 2x the actual data length, triggering:
    
      TCP recvmsg seq # bug 2: copied BF2E806, seq BF2E7FD, \
                               rcvnxt BF2E806, fl 0
      WARNING: net/ipv4/tcp.c:2745 at tcp_recvmsg_locked+0x72b/0x2640
      Call Trace:
       tcp_recvmsg+0x10a/0x500
       sock_recvmsg+0x168/0x1d0
       __sys_recvfrom+0x19a/0x2a0
       __x64_sys_recvfrom+0xe4/0x1f0
       do_syscall_64+0xf7/0x530
       entry_SYSCALL_64_after_hwframe+0x77/0x7f
    
      cleanup rbuf bug: copied BF2E806 seq BF2E806 rcvnxt BF2E806
      WARNING: net/ipv4/tcp.c:1609 at tcp_cleanup_rbuf+0xf2/0x1c0
      Call Trace:
       tcp_recvmsg_locked+0x8d1/0x2640
       tcp_recvmsg+0x10a/0x500
       sock_recvmsg+0x168/0x1d0
       __sys_recvfrom+0x19a/0x2a0
       __x64_sys_recvfrom+0xe4/0x1f0
       do_syscall_64+0xf7/0x530
       entry_SYSCALL_64_after_hwframe+0x77/0x7f
    
    Fix this by converting self-redirect verdict to __SK_PASS at the
    beginning of sk_psock_verdict_apply(). This bypasses the
    __SK_REDIRECT case entirely (which calls sk_psock_eat_skb), letting
    the __SK_PASS path queue the skb to the psock ingress queue. The
    data is then read via tcp_bpf_recvmsg_parser(), which advances
    copied_seq exactly once through copied_from_self. Cross-socket
    redirects continue through __SK_REDIRECT with sk_psock_eat_skb()
    unchanged.
    
    Fixes: e5c6de5fa025 ("bpf, sockmap: Incorrectly handling copied_seq")
    Suggested-by: Jakub Sitnicki <[email protected]>
    Suggested-by: Jiayuan Chen <[email protected]>
    Signed-off-by: Geliang Tang <[email protected]>
    Reviewed-by: Emil Tsalapatis <[email protected]>
    Reviewed-by: Jiayuan Chen <[email protected]>
    Link: https://lore.kernel.org/r/1a8e797a1b26e2f695aaac22ac644c2862f63466.1788858299.git.tanggeliang@kylinos.cn
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc [+ + +]
Author: Zhao Gongyi <[email protected]>
Date:   Thu Sep 17 20:10:16 2026 +0800

    bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc
    
    [ Upstream commit 814a81c842bd88f6bd8a4ce550d560df071a5d03 ]
    
    sock_map_alloc() only rejects max_entries == 0 and otherwise allows any
    u32 value.  sock_map_free() then walks the sks[] array with a signed int
    iterator:
    
            int i;
            for (i = 0; i < stab->map.max_entries; i++)
                    struct sock **psk = &stab->sks[i];
    
    When a SOCKMAP is created with max_entries = 0xffffffff (UINT_MAX), the
    allocation of 32 GiB can succeed on large-memory hosts.  During free the
    counter reaches 0x80000000, wraps to INT_MIN, is sign-extended by movslq
    and turned into a ~16 GiB negative offset from stab->sks, pointing far
    below the allocation.
    
    The faulting access is an xchg() write in sock_map_free().  Without
    KASAN, the same out-of-bounds write can fault on an unmapped vmalloc page
    or corrupt an unrelated allocation if that vmalloc address is populated.
    On a KASAN kernel with CONFIG_KASAN_VMALLOC=y, the shadow check for that
    address hits an unmapped shadow page and oopses first:
    
      BUG: unable to handle page fault for address: fffff521b59c5a00
      RIP: 0010:kasan_check_range+0x107/0x190
      Call Trace:
       sock_map_free+0x93/0x190
       map_create+0x68d/0xb30
       __sys_bpf+0x21e/0x2e70
    
    Vmcore confirmed stab->map.max_entries == 0xffffffff, stab->sks ==
    0xffffc911ace2d000, and the faulting address sks + (s64)INT_MIN * 8
    exactly at 0xffffc90dace2d000.  The same buggy path is reached on the
    normal close()/bpf_map_free_deferred() path whenever such a map is
    destroyed.
    
    sock_map_alloc() used to bound its allocation size through
    bpf_map_charge_init(), but the bound was dropped when rlimit-based memory
    accounting was removed.  Reject max_entries > INT_MAX at creation time so
    the signed iterator in sock_map_free() never sees a value that would
    overflow.
    
    Triggered by syzkaller and reproduced on both a 6.6-based KASAN kernel
    and the upstream v7.3-rc2 kernel.
    
    Fixes: 0d2c4f964050 ("bpf: Eliminate rlimit-based memory accounting for sockmap and sockhash maps")
    Signed-off-by: Zhao Gongyi <[email protected]>
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

 
bpf: Allow terminal gotox instructions [+ + +]
Author: Siddharth Chintamaneni <[email protected]>
Date:   Wed Sep 2 17:14:13 2026 +0000

    bpf: Allow terminal gotox instructions
    
    [ Upstream commit 0d7823cd4cda35f2060a685fa276ff7709915abc ]
    
    check_subprogs() treats gotox as a direct jump and validates its reserved
    zero offset. When gotox is the final instruction, this produces a
    synthetic successor one instruction past the end of the subprogram and
    rejects an otherwise valid program.
    
    Skip direct-offset validation for gotox and accept it as a
    non-fallthrough terminal instruction. Its actual targets remain validated
    from the instruction-array jump table during CFG construction.
    
    Fixes: 493d9e0d6083 ("bpf, x86: add support for indirect jumps")
    Signed-off-by: Siddharth Chintamaneni <[email protected]>
    Reviewed-by: Anton Protopopov <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Avoid soft lockup in __htab_map_lookup_and_delete_batch() [+ + +]
Author: Jose Fernandez (Anthropic) <[email protected]>
Date:   Wed Sep 9 17:51:04 2026 +0000

    bpf: Avoid soft lockup in __htab_map_lookup_and_delete_batch()
    
    [ Upstream commit 85136bf22404474a815fc0ed26ec0d1cbc1bc3f9 ]
    
    __htab_map_lookup_and_delete_batch() has no rescheduling point. The
    batch count bounds how many entries are copied out, not how many
    buckets are visited, so one BPF_MAP_LOOKUP_BATCH call can walk the
    map end to end. The empty-bucket fast path is worse: it stays inside
    a single rcu_read_lock() / bpf_disable_instrumentation() section for
    any run of consecutive empty buckets.
    
    That holds up on small maps, but it falls apart at scale. On a
    144-CPU arm64 host running a CONFIG_PREEMPT_NONE kernel, periodic
    BPF_MAP_LOOKUP_BATCH calls against an LRU hash map with 16,777,216
    buckets held a CPU inside the batch op for 77+ seconds and triggered
    the soft lockup watchdog.
    
    Commit 75134f16e7dd ("bpf: Add schedule points in batch ops") fixed this
    same problem in the generic batch ops, but not in this htab-native path,
    which every htab-based hash map variant uses for its lookup[_and_delete]
    batch ops.
    
    Complete that fix here. Leave the critical section after 64 consecutive
    empty buckets, call cond_resched_tasks_rcu_qs(), and resume at the saved
    bucket cursor. No locks are held at that point, and resuming from the
    cursor is already the function's behavior for non-empty buckets. Add the
    same call to the per-bucket loop after copy_to_user(), where every lock
    has been dropped. cond_resched_rcu() is not enough here: sleeping with
    bpf_prog_active elevated makes tracing programs on that CPU silently
    skip their invocations.
    
    Plain cond_resched() is not enough either. It is a no-op under PREEMPT
    and PREEMPT_LAZY, the only models arm64 and x86 have offered since
    commit 7dadeaa6e851 ("sched: Further restrict the preemption modes").
    It is also never a Tasks RCU quiescent state, in any model: the
    reschedule counts as a preemption. The walking task stays a holdout and
    stalls every synchronize_rcu_tasks() caller, ftrace and BPF trampoline
    teardown included, until the syscall returns [1].
    cond_resched_tasks_rcu_qs() is the usual tool for that [2]. It reports
    the quiescent state at each yield and still reschedules as
    cond_resched() does on PREEMPT_NONE and PREEMPT_VOLUNTARY kernels.
    
    Fixes: 057996380a42 ("bpf: Add batch ops to all htab bpf map")
    Cc: "Paul E. McKenney" <[email protected]>
    Cc: Rik van Riel <[email protected]>
    Link: https://lore.kernel.org/bpf/20260715215314.44423f47@fangorn/ [1]
    Link: https://lore.kernel.org/bpf/9d444098-7c03-4163-af12-bd0a79a51443@paulmck-laptop/ [2]
    Assisted-by: LLM
    Signed-off-by: Jose Fernandez (Anthropic) <[email protected]>
    Signed-off-by: Josef Bacik <[email protected]>
    Reviewed-by: Rik van Riel <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Bound ownership depth through local kptrs and graph roots [+ + +]
Author: Kumar Kartikeya Dwivedi <[email protected]>
Date:   Mon Sep 14 15:24:42 2026 +0200

    bpf: Bound ownership depth through local kptrs and graph roots
    
    [ Upstream commit bfc888f04588f591851e95c974954cfca58e6c19 ]
    
    Program-allocated objects can own other local objects through referenced
    kptrs. bpf_obj_free_fields() follows those pointers through
    __bpf_obj_drop_impl() synchronously, before the object storage is freed
    through RCU. A self-referential local kptr type therefore permits arbitrarily
    deep object chains, and dropping the head can exhaust the kernel stack.
    Long acyclic type chains have the same problem.
    
    btf_check_and_fixup_fields() still assumes referenced kptrs only point to
    kernel types and checks ownership through list and rbtree roots only. Its
    existing rule is sufficient for graph-only cycles: the target of each graph
    edge must contain a node, so every type in a cycle has both a root and a
    node. The rule rejects such a type owning another root, breaking every
    cycle. It also limits graph-only chains to three types, or two if the first
    type contains a node, and conservatively rejects longer acyclic chains.
    The missing local-kptr edges, rather than a missed graph-only cycle, are the
    bug introduced by support for bpf_kptr_xchg() into local kptrs.
    
    Replace that restriction with one bounded ownership walk covering graph
    roots and local referenced kptrs. Run it after all BTF records have been
    fixed up, reject cycles and paths deeper than eight record-bearing types,
    and cache each type's suffix depth while checking it against the remaining
    budget. This also permits the longer acyclic graph-only layouts rejected
    by the old rule; update their existing BTF tests accordingly.
    
    Keep the bound independent of MAX_CALL_FRAMES because recursive destruction
    can run below a BPF call chain. A plain local pointee without special-field
    metadata adds only a final non-recursing drop. Non-owning kptrs and
    kernel-BTF kptrs do not recurse through local records and remain outside the
    walk. Include local percpu-kptr edges too, although allocation of percpu
    objects with special fields is currently forbidden, so that relaxing that
    restriction cannot bypass the ownership bound.
    
    btf_check_and_fixup_fields() continues to initialize graph_root.value_rec,
    including for separately allocated map records. The ownership relationships
    belong to immutable program BTF and only need validation at BTF load time.
    
    Fixes: b0966c724584 ("bpf: Support bpf_kptr_xchg into local kptr")
    Reported-by: Nicholas Carlini <[email protected]>
    Suggested-by: Nicholas Carlini <[email protected]>
    Signed-off-by: Kumar Kartikeya Dwivedi <[email protected]>
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Check params size before reading reserved fields [+ + +]
Author: Yuqi Xu <[email protected]>
Date:   Sat Sep 19 16:45:02 2026 +0800

    bpf: Check params size before reading reserved fields
    
    [ Upstream commit a11212910cf09b2fe8db9afa41ef60c4f81879c5 ]
    
    bpf_crypto_ctx_create() is a kfunc whose second argument is declared
    with the __sz annotation, so the verifier only guarantees that
    params__sz bytes of params are valid.  The function nevertheless reads
    params->reserved[0] and params->reserved[1] (offsets 14 and 15) before
    comparing params__sz against the size of struct bpf_crypto_params, so a
    BPF program can pass a shorter buffer and have the kernel read past the
    region that was validated for it.
    
    Move the size check in front of the reserved field reads.
    
    Fixes: 3e1c6f35409f ("bpf: make common crypto API for TC/XDP programs")
    Reported-by: Vega <[email protected]>
    Signed-off-by: Yuqi Xu <[email protected]>
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Reviewed-by: Ren Wei <[email protected]>
    Link: https://patch.msgid.link/4f3ab4b03e79017e215521743996555439bf0bb3.1789802413.git.xuyuqiabc@gmail.com
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL [+ + +]
Author: Weiming Shi <[email protected]>
Date:   Wed Sep 9 12:08:08 2026 +0800

    bpf: Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL
    
    [ Upstream commit e4a62833adff6ef0fe7c0b90393204fe3c26b5c5 ]
    
    An LWT_SEG6LOCAL program can invalidate its cached SRH with
    bpf_lwt_seg6_adjust_srh() and then call bpf_skb_pull_data(). The latter
    may reallocate skb->head, leaving the per-CPU SRH pointer dangling.
    Post-program SRH validation then writes through that pointer.
    
    Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL programs so the verifier
    rejects this unsafe helper combination. Other LWT program types continue
    to expose the helper through lwt_out_func_proto().
    
    Fixes: 004d4b274e2a ("ipv6: sr: Add seg6local action End.BPF")
    Reported-by: [email protected]
    Suggested-by: Alexei Starovoitov <[email protected]>
    Signed-off-by: Weiming Shi <[email protected]>
    Signed-off-by: Daniel Borkmann <[email protected]>
    Reviewed-by: Emil Tsalapatis <[email protected]>
    Closes: https://lore.kernel.org/all/[email protected]/
    Link: https://lore.kernel.org/bpf/[email protected]/
    Link: https://lore.kernel.org/bpf/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Fix bounds check for skb-backed dynptrs [+ + +]
Author: Emil Tsalapatis <[email protected]>
Date:   Tue Sep 22 17:20:18 2026 +0000

    bpf: Fix bounds check for skb-backed dynptrs
    
    [ Upstream commit ed6eec97b534979dcf28b40c389cee57bd6561d4 ]
    
    The skb_pointer_if_linear() function checks whether a
    memory region of length len starting at offset off into
    the skb is in the linear area, and returns a pointer to
    the region if so. The check currently subtracts between
    skb_headlen and offset of the check, and since skb_headlen
    is unsigned the subtraction can underflow. This causes the
    bounds check to spuriously pass and generate an arbitrary
    pointer of the form *(skb->data + off).
    
    The only user of this helper is currently skb-backed BPF
    dynptr code. Returning the wrong pointer leads to the
    dynptr erroneously being backed with invalid memory.
    
    Ensure the subtraction cannot underflow, and fail the check if
    it would. Use u64 arithmetic to also prevent overflow when
    calculating (skb_headlen(skb) - off) since off is unsigned.
    
    Fixes: 6f5a630d7c57 ("bpf, net: Introduce skb_pointer_if_linear().")
    Reported-by: Nicholas Carlini <[email protected]>
    Signed-off-by: Emil Tsalapatis <[email protected]>
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Reviewed-by: Jiayuan Chen <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Fix bpf_skb_change_tail wrt csum partial skbs [+ + +]
Author: Daniel Borkmann <[email protected]>
Date:   Mon Sep 7 14:10:24 2026 +0200

    bpf: Fix bpf_skb_change_tail wrt csum partial skbs
    
    [ Upstream commit 3b55f350c68a0aceff108f47f9d31f47ebffaf7b ]
    
    Cilium generates ICMP "frag needed" replies from BPF when a LB DSR
    packet exceeds the egress MTU. The reply is built by first trimming the
    packet down to target size via bpf_skb_change_tail(), and then pushing
    the ICMP error headers in front of it.
    
    The trim is rejected for skbs which carry a checksum offload, e.g. TCP
    packets aggregated by GRO on ingress where tcp_gro_complete() leaves
    the skb as CHECKSUM_PARTIAL. __bpf_skb_min_len() raises the minimum
    length to the end of the L4 checksum field, so a trim to 42 bytes bails
    out with -EINVAL given a min_len of 52 in this case, and due to that
    the ICMP generator fails. This is not the case if GRO is turned off.
    
    Fix this bpf_skb_change_tail() restriction and drop the checksum offload
    when the new length no longer covers the checksum field. The BPF program
    rewrites the skb into an ICMP error and computes the checksum itself
    anyway.
    
    Fixes: 5293efe62df8 ("bpf: add bpf_skb_change_tail helper")
    Reported-by: Tom Hadlaw <[email protected]>
    Reported-by: Yusuke Suzuki <[email protected]>
    Signed-off-by: Daniel Borkmann <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Fix bpf_sock context code generation [+ + +]
Author: Emil Tsalapatis <[email protected]>
Date:   Tue Sep 22 17:20:20 2026 +0000

    bpf: Fix bpf_sock context code generation
    
    [ Upstream commit 4a4852376e3a2727ea40e61143d6d7c22bb6dfad ]
    
    Currently, the ctx access code reads the rx_queue_mapping
    field with either a 4-byte or 2-byte load. The rest of the bits
    in the register are marked known zero by the verifier. However,
    the emitted ctx access code places in the register on certain
    the special value (-1) using BPF_MOV_IMM64, which gets sign-extended
    to turn on all the bits in the register. By shifting this value right,
    the program ends up with a value at runtime above what the verifier
    assumes is possible.
    
    Fix this by ensuring the read value is as wide as the assumed size.
    Use MOV32 instructions instead of MOV64 instructions to keep
    the upper bits zero as assumed by the verifier. Also properly report
    the size of the destination variable (the bpf_sock field, 4 bytes) instead
    of the source (the socket field, 2 bytes).
    
    Fixes: c3c16f2ea6d2 ("bpf: Add rx_queue_mapping to bpf_sock")
    Reported-by: Nicholas Carlini <[email protected]>
    Suggested-by: Nicholas Carlini <[email protected]>
    Signed-off-by: Emil Tsalapatis <[email protected]>
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Reviewed-by: Jiayuan Chen <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Fix BSWAP 32 and 16 on MIPS64 [+ + +]
Author: Johan Almbladh <[email protected]>
Date:   Wed Sep 23 12:51:58 2026 +0200

    bpf: Fix BSWAP 32 and 16 on MIPS64
    
    [ Upstream commit 8110ba09777873443db286b3cbb89b0e6311c554 ]
    
    The 16/32-bit byteswap implementations for MIPS64r1 and earlier do
    not have an explicit zero extension afterwards. The input is first
    sign-extended to 64 bits, and the byteswap sequence can then leave
    the result sign-extended depending on the value of the low bits.
    
    Add the missing zero-extension.
    
    Found with test_bpf on MIPS64r1 emulated by QEMU.
    
    Fixes: fbc802de6b10 ("mips, bpf: Add new eBPF JIT for 64-bit MIPS")
    Signed-off-by: Johan Almbladh <[email protected]>
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Fix divide-by-zero in btf_struct_walk() [+ + +]
Author: Jiayuan Chen <[email protected]>
Date:   Thu Sep 10 20:22:55 2026 +0800

    bpf: Fix divide-by-zero in btf_struct_walk()
    
    [ Upstream commit b0b3dc66529676228cb938cbcad66920f735c223 ]
    
    When an access goes past the struct and the last member is a flexible
    array, btf_struct_walk() folds the offset back into a single element with
    (off - moff) % t->size, but never checks that the element type has a size.
    
    BTF takes an empty struct, so this in program BTF
    
            /* event could be empty */
            struct event {
     #ifdef HAVE_TIMESTAMP
                    __u64 ts;
     #endif
            };
    
            struct batch {
                    int nr;
                    struct event events[];
            };
    
    divides by zero at prog load time. Getting there needs a PTR_TO_BTF_ID that
    is not MEM_ALLOC, e.g. a plain read of a local kptr stashed in a map from a
    sleepable program.
    
    Oops: divide error: 0000 [#1] SMP KASAN PTI
    RIP: 0010:btf_struct_walk+0x53f/0x1570
    Call Trace:
     <TASK>
     btf_struct_access+0x42a/0xcd0
     check_ptr_to_btf_access+0x4dc/0x1160
     check_mem_access+0x3a45/0x8740
     check_load_mem+0x36a/0xd10
     do_check_common+0x3ef0/0xb210
     bpf_check+0x6d3b/0x8580
     bpf_prog_load+0xf7c/0x2720
     __sys_bpf+0xa83/0x3690
     __x64_sys_bpf+0xc7/0x150
     x64_sys_call+0x1f3f/0x27e0
     do_syscall_64+0xe5/0x610
     entry_SYSCALL_64_after_hwframe+0x76/0x7e
     </TASK>
    
    Reject a zero-sized element type. The fixed array path in the same function
    already bails out on the same thing:
    
            btf_struct_walk()
            ...
                    /* skip empty array */
                    if (moff == mtrue_end)
                            continue;
    
                    msize /= total_nelems;
    
    Fixes: 9c5f8a1008a1 ("bpf: Support variable length array in tracing programs")
    Signed-off-by: Jiayuan Chen <[email protected]>
    Acked-by: Eduard Zingerman <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Fix immediate JMP JEQ/JNE on MIPS32 [+ + +]
Author: Johan Almbladh <[email protected]>
Date:   Wed Sep 23 12:51:57 2026 +0200

    bpf: Fix immediate JMP JEQ/JNE on MIPS32
    
    [ Upstream commit db762fd96be225bd06161c9631160c755d891693 ]
    
    An addu instruction was emitted instead of addiu, causing the immediate
    value 1 to be interpreted as register $at. This made the comparison
    result invalid when the immediate operand was negative. Note that $at
    is mapped to BPF_REG_AX, which is used for constant blinding.
    
    Fix the instruction to use the immediate form.
    
    Found with test_bpf on MIPS32r1 emulated by QEMU.
    
    Fixes: eb63cfcd2ee8 ("mips, bpf: Add eBPF JIT for 32-bit MIPS")
    Signed-off-by: Johan Almbladh <[email protected]>
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Fix out-of-bounds read of rtt_min in sock_ops [+ + +]
Author: Jiayuan Chen <[email protected]>
Date:   Thu Sep 3 18:09:20 2026 +0800

    bpf: Fix out-of-bounds read of rtt_min in sock_ops
    
    [ Upstream commit 75f8cf22463d82bb1fb0239a3d485fc8f4c8ef03 ]
    
    A sockops prog reading skops->rtt_min never checks the sk type: on the
    tcp_conn_request() path sock_ops->sk is a request_sock (non-full), and the
    ctx rewrite casts it to a tcp_sock (full) and reads rtt_min past the end of
    the request_sock, returning dirty adjacent memory.
    
            SEC("sockops")
            int prog(struct bpf_sock_ops *skops)
            {
                    switch (skops->op) {
                    case BPF_SOCK_OPS_RWND_INIT:
                            leak = skops->rtt_min;   /* reads the request_sock OOB */
                    ...
                    }
            }
    
    For instance one such read returned rtt_min=0xffff8881, the high half of a
    leaked kernel pointer.
    
    Guarding that cast is exactly what SOCK_OPS_GET_FIELD() does -- it checks
    is_locked_tcp_sock and returns 0 when sock_ops->sk is not a locked full
    socket. Every other tcp_sock field in sock_ops goes through it; rtt_min is
    the only one open-coded, so it skips the check.
    
    Read rtt_min through SOCK_OPS_GET_FIELD() too. rtt_min is a bit special:
    it is a struct minmax and we only want the current min, so pass
    rtt_min.s[0].v. That is equivalent to the old hand-computed offset
    
            offsetof(struct tcp_sock, rtt_min) + sizeof_field(struct minmax_sample, t)
    
    (s[0] sits at rtt_min + 0 and .v at + sizeof(.t), i.e. what minmax_get()
    returns), so the loaded field is unchanged and only the full-sock guard is
    added. The two BUILD_BUG_ON()s that protected the hand-computed offset
    are no longer needed.
    
    Before patch:
    
            0: r1 = *(u64 *)(r1 +0)      ; r1 = skops->sk
            1: r1 = *(u32 *)(r1 +2324)   ; ((tcp_sock *)sk)->rtt_min.s[0].v
    
    After patch:
    
            0: *(u64 *)(r1 +56) = r9
            1: r9 = *(u8 *)(r1 +50)      ; is_locked_tcp_sock
            2: if r9 == 0 goto pc+4      ; not a locked full sock -> 0
            3: r9 = *(u64 *)(r1 +56)
            4: r1 = *(u64 *)(r1 +0)      ; r1 = skops->sk
            5: r1 = *(u32 *)(r1 +2324)   ; rtt_min.s[0].v
            6: goto pc+2
            7: r9 = *(u64 *)(r1 +56)
            8: r1 = 0
    
    Fixes: 44f0e43037d3 ("bpf: Add support for reading sk_state and more")
    Reported-by: VEGA <[email protected]>
    Signed-off-by: Jiayuan Chen <[email protected]>
    Reviewed-by: Emil Tsalapatis <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Fix out-of-bounds read of sk_protocol in bpf_sock_destroy() [+ + +]
Author: Jiayuan Chen <[email protected]>
Date:   Thu Sep 10 19:26:26 2026 +0800

    bpf: Fix out-of-bounds read of sk_protocol in bpf_sock_destroy()
    
    [ Upstream commit 01b245ba016d44861690594e10f67e026ce8552f ]
    
    sk_protocol lives in struct sock, not in struct sock_common. A timewait
    or request sock handed to bpf_sock_destroy() by the tcp iterator is
    neither, so reading sk->sk_protocol runs past the object:
    
    ==================================================================
    BUG: KASAN: slab-out-of-bounds in bpf_sock_destroy+0xc7/0xe0
    Read of size 2 at addr ffff8881047d11b4 by task test_progs/428
    
    Tainted: [W]=WARN
    Call Trace:
     <TASK>
     dump_stack_lvl+0x91/0xf0
     print_report+0xd1/0x630
     kasan_report+0xf3/0x130
     __asan_report_load2_noabort+0x14/0x30
     bpf_sock_destroy+0xc7/0xe0
     bpf_prog_c3dd61f9d9cd9f37_iter_tcp6_timewait+0x9f/0xb7
     bpf_iter_run_prog+0x538/0xde0
     bpf_iter_tcp_seq_show+0x26b/0x4b0
     bpf_seq_read+0x424/0x1210
     vfs_read+0x197/0xe40
     ksys_read+0x119/0x240
     __x64_sys_read+0x72/0xc0
     x64_sys_call+0x647/0x27e0
     do_syscall_64+0xe5/0x610
     entry_SYSCALL_64_after_hwframe+0x76/0x7e
    
    Only check sk_protocol on full socks. tcp_abort() already knows how to
    deal with TIME_WAIT and NEW_SYN_RECV socks. Also fix the comment, it
    never matched the code.
    
    Fixes: 4ddbcb886268 ("bpf: Add bpf_sock_destroy kfunc")
    Reported-by: Xiang Mei (Microsoft) <[email protected]>
    Closes: https://lore.kernel.org/bpf/[email protected]/
    Signed-off-by: Jiayuan Chen <[email protected]>
    Reviewed-by: Kuniyuki Iwashima <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Fix u32 overflow issue in map batch operations [+ + +]
Author: Masoud Aghasi <[email protected]>
Date:   Thu Sep 3 09:27:34 2026 +0100

    bpf: Fix u32 overflow issue in map batch operations
    
    [ Upstream commit 953824e508b27d12837e32ef37ef6248e1f6fc7a ]
    
    Several map batch operation implementations such as
    generic_map_lookup_batch() use calculations in the form of
    "values + cp * map->value_size" to compute the desired userspace memory
    address for reading or writing. This can overflow the u32 type
    (the result of "cp * map->value_size") when the map size exceeds 4GB.
    
    generic_map_lookup_batch() may corrupt values for some keys in
    userspace memory, and in some cases it mismatches values for some keys
    while still reporting success.
    
    Other batch operations may fail to delete or update some keys,
    or the syscall may return unexpected errors.
    
    Add size_t casts to prevent the affected offset and size calculations
    from overflowing.
    
    Fixes: cb4d03ab499d ("bpf: Add generic support for lookup batch op")
    Fixes: aa2e93b8e58e ("bpf: Add generic support for update and delete batch ops")
    Fixes: 057996380a42 ("bpf: Add batch ops to all htab bpf map")
    Signed-off-by: Masoud Aghasi <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Fix UAF due to concurrent consumption of ttrace lists in alloc_bulk [+ + +]
Author: Pu Lehui <[email protected]>
Date:   Sat Sep 5 02:11:39 2026 +0000

    bpf: Fix UAF due to concurrent consumption of ttrace lists in alloc_bulk
    
    [ Upstream commit 1c21452d02eec2f008e2c5535820f85adbd7587a ]
    
    Syzkaller repeatedly triggered UAF splats related to nodes in
    waiting_for_gp_ttrace within the bpf memalloc:
    
    BUG: KASAN: slab-use-after-free in llist_del_first+0x85/0x110 lib/llist.c:61
    Read of size 8 at addr ffff8881572cd080 by task syz.4.470/5112
     ...
     llist_del_first+0x85/0x110 lib/llist.c:61
     alloc_bulk+0x193/0x460 kernel/bpf/memalloc.c:229
     bpf_mem_refill+0x386/0x560 kernel/bpf/memalloc.c:436
    
    Freed by task 14:
     ...
     __free_rcu kernel/bpf/memalloc.c:281 [inline]
     __free_rcu_tasks_trace+0x48/0xd0 kernel/bpf/memalloc.c:291
     rcu_tasks_invoke_cbs+0x1ec/0x3e0 kernel/rcu/tasks.h:571
     rcu_tasks_one_gp+0x13d/0x220 kernel/rcu/tasks.h:621
     rcu_tasks_kthread+0xf3/0x120 kernel/rcu/tasks.h:651
    
    The reason is that the UAF occurs after the RCU Tasks Trace GP expires:
    when the __free_rcu() callback runs, there is no synchronization
    protecting llist_del_all() against concurrent alloc_bulk() operating on
    waiting_for_gp_ttrace, leading to the race condition below:
    
    CPU0                                           CPU1
                                                   __free_rcu (RCU Tasks Trace callback)
    alloc_bulk
      llist_del_first(&c->waiting_for_gp_ttrace)
        entry = smp_load_acquire(&head->first);
        do {
          if (entry == NULL)
            return NULL;
                                                   free_all(llist_del_all(&c->waiting_for_gp_ttrace))
                                                     llist_for_each_safe(pos, t, llnode)
                                                       free_one(pos);
          next = READ_ONCE(entry->next); <-- trigger UAF
        } while (!try_cmpxchg(&head->first, &entry, next));
    
    In addition, there is also a theoretical race condition on the
    free_by_rcu_ttrace list. This race requires two preconditions: an
    in-flight Tasks Trace GP keeping c->call_rcu_ttrace_in_progress == 1,
    and concurrent cross-CPU frees repopulating c->free_by_rcu_ttrace with
    new nodes. Under these conditions, the following scenario triggers UAF:
    
    // CPU0
    // irq work is still busy (on PREEMPT_RT)
    alloc_bulk()
      llist_del_first(&c->free_by_rcu_ttrace)
        entry = smp_load_acquire(&head->first);
        do {
          if (entry == NULL)
            return NULL;
    
            // CPU1
            bpf_mem_alloc_destroy()
              WRITE_ONCE(c->draining, true)
              // wait for CPU0
              irq_work_sync()
    
                    // CPU2
                    do_call_rcu_ttrace(tgt(CPU0))
                      if (c->draining) {
                        llist_del_all(&c->free_by_rcu_ttrace)
                        free_all()
                      }
    
    // CPU0 continue
          next = READ_ONCE(entry->next); <-- trigger UAF
        while (!try_cmpxchg(&head->first, &entry, next));
    
    Fix this by introducing a raw spinlock to synchronize the concurrent
    consumption on waiting_for_gp_ttrace and free_by_rcu_ttrace.
    
    Fixes: 04fabf00b4d3 ("bpf: Allow reuse from waiting_for_gp_ttrace list.")
    Suggested-by: Alexei Starovoitov <[email protected]>
    Suggested-by: Hou Tao <[email protected]>
    Signed-off-by: Pu Lehui <[email protected]>
    Acked-by: Hou Tao <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

bpf: fs/xattr: don't assume the inode is locked in path_unlink/path_rmdir [+ + +]
Author: Andrea Parri <[email protected]>
Date:   Tue Sep 22 16:55:30 2026 +0200

    bpf: fs/xattr: don't assume the inode is locked in path_unlink/path_rmdir
    
    commit 35d442ed1f86465e49df3119fb898f985186db13 upstream.
    
    bpf_lsm_has_d_inode_locked() makes the verifier rewrite
    bpf_[set|remove]_dentry_xattr() to the _locked variants, which assume
    that the caller already holds the inode's i_rwsem.  The path_unlink and
    path_rmdir hooks are listed, but security_path_unlink() and
    security_path_rmdir() run before vfs_unlink()/vfs_rmdir() take the
    victim inode's i_rwsem, so a sleepable BPF LSM program attached to
    either hook mutates the victim's xattrs without the lock held.
    
    Drop the two path hooks from d_inode_locked_hooks so that the verifier
    keeps the locking bpf_[set|remove]_dentry_xattr() variants, which take
    the lock themselves.
    
    Fixes: 56467292794b8 ("bpf: fs/xattr: Add BPF kfuncs to set and remove xattrs")
    Cc: [email protected]
    Signed-off-by: Andrea Parri <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Christian Brauner (Amutable) <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

bpf: Make post-verification instruction rewrites killable [+ + +]
Author: Kumar Kartikeya Dwivedi <[email protected]>
Date:   Fri Sep 18 01:32:09 2026 +0200

    bpf: Make post-verification instruction rewrites killable
    
    [ Upstream commit 261b61d3735b042ae25634f795c4540be0fc140c ]
    
    After do_check() returns, the verifier runs several instruction rewrite
    passes. Some of them patch or remove one instruction at a time. Each
    operation moves the remaining instruction and auxiliary-data arrays and
    adjusts all branch offsets, making the overall work quadratic in the
    program length.
    
    A privileged loader can submit 131072 unconditional jumps by zero followed
    by a valid return. Verification finishes quickly, but bpf_opt_remove_nops()
    then spends a long time removing each jump separately. Since this
    post-verification work neither checks for signals nor reschedules, a pending
    SIGKILL cannot terminate the task until the rewrite finishes.
    
    Make bpf_patch_insn_data() and verifier_remove_insns() common cancellation
    and rescheduling points. These helpers run from BPF_PROG_LOAD process
    context, and bpf_patch_insn_data() can already sleep while reallocating
    auxiliary data.
    
    Report interrupted constant blinding as -EINTR and propagate it through
    both JIT paths, including kernels that permit interpreter fallback.
    Other blinding failures retain the existing fallback behavior.
    
    This does not reduce the quadratic cost of the rewrite passes, but it makes
    the work preemptible and allows a killed loader to be torn down promptly.
    
    Fixes: 52875a04f4b2 ("bpf: verifier: remove dead code")
    Reported-by: Nicholas Carlini <[email protected]>
    Suggested-by: Nicholas Carlini <[email protected]>
    Signed-off-by: Kumar Kartikeya Dwivedi <[email protected]>
    Acked-by: Eduard Zingerman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Eduard Zingerman <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Preserve packet pointer class displacement in regsafe() [+ + +]
Author: Kumar Kartikeya Dwivedi <[email protected]>
Date:   Fri Sep 18 01:32:10 2026 +0200

    bpf: Preserve packet pointer class displacement in regsafe()
    
    [ Upstream commit fd16449a9b3b31a8f18944c2f0e29e4e218ca2cf ]
    
    regsafe() maps packet pointer IDs between states and checks that each
    current register range is a subset of the corresponding explored
    register range. It does not, however, preserve the displacement between
    registers that share a packet pointer ID.
    
    This is unsound because packet range is shared by ID. A bounds check on
    one class member updates every member, and a later access can consume the
    range through another member. Commit 022ac0750883 ("bpf: use reg->var_off
    instead of reg->off for pointers") folded the fixed pointer offset into
    r64 and removed the old off equality check, so two individually narrower
    registers can prune even when their displacement has changed. The
    explored path can then license an out-of-bounds packet access on the
    pruned path.
    
    Require matching range bases for packet pointers with an ID. Together
    with the existing ID mapping, this preserves the displacement between
    members of each packet-pointer class without adding per-ID state.
    Packet pointers without an ID remain unaffected.
    
    Fixes: 022ac0750883 ("bpf: use reg->var_off instead of reg->off for pointers")
    Reported-by: Nicholas Carlini <[email protected]>
    Suggested-by: Nicholas Carlini <[email protected]>
    Signed-off-by: Kumar Kartikeya Dwivedi <[email protected]>
    Acked-by: Eduard Zingerman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Eduard Zingerman <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Reject dev-bound-only programs on other devices [+ + +]
Author: Weiming Shi <[email protected]>
Date:   Sun Sep 20 21:23:04 2026 +0800

    bpf: Reject dev-bound-only programs on other devices
    
    [ Upstream commit 6db1ce73e9853f533eb7f413f14ba00f8ec6f80d ]
    
    __bpf_offload_dev_match() falls back to comparing offdev pointers after an
    exact netdev mismatch. Bound-only programs normally have NULL offdevs, so
    unrelated netdevs compare equal. A bound-only program on an
    offload-registered netdev can instead inherit a real offdev and match a
    sibling port. With CAP_BPF and CAP_NET_ADMIN, a caller can use
    bpf(BPF_LINK_CREATE) with a different target ifindex to run metadata kfuncs
    specialized for the bound driver on the target driver's xdp_buff. Running a
    veth-bound program on tun reads beyond tun's bare stack xdp_buff as a
    veth_xdp_buff.
    
      Oops: general protection fault, probably for non-canonical address
      KASAN: null-ptr-deref in range [0x0000000000000010-0x0000000000000017]
      RIP: 0010:veth_xdp_rx_timestamp (drivers/net/veth.c:1673)
      Call Trace:
       ...
       tun_build_skb (drivers/net/tun.c:1739)
       tun_get_user (drivers/net/tun.c:1856)
       tun_chr_write_iter (drivers/net/tun.c:2091)
       vfs_write (fs/read_write.c:595 fs/read_write.c:687)
       ksys_write (fs/read_write.c:739)
       do_syscall_64 (arch/x86/entry/syscall_64.c:84)
       entry_SYSCALL_64_after_hwframe (arch/x86/entry/entry_64.S:121)
      Kernel panic - not syncing: Fatal exception in interrupt
    
    Restrict non-offloaded programs to exact netdev matches and retain the
    shared-offdev fallback only for genuinely offloaded multi-port programs.
    
    Fixes: 2b3486bc2d23 ("bpf: Introduce device-bound XDP programs")
    Reported-by: <[email protected]>
    Signed-off-by: Weiming Shi <[email protected]>
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Link: https://lore.kernel.org/bpf/[email protected]/
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Reject non-negative offsets in stack_slot_obj_get_spi() [+ + +]
Author: Xu Yunxiang <[email protected]>
Date:   Mon Sep 21 05:04:21 2026 +0800

    bpf: Reject non-negative offsets in stack_slot_obj_get_spi()
    
    [ Upstream commit 79a9172f3ab4ad8392c5e5c8944b9b7710ade620 ]
    
    bpf_get_spi() computes (-off - 1) / BPF_REG_SIZE using C division,
    which truncates toward zero. For off == 0, this produces spi 0, the
    same index used by the valid stack slot at fp-8.
    
    stack_slot_obj_get_spi() currently checks alignment and the resulting
    spi bounds, but does not reject the non-negative offset itself. It can
    therefore validate a PTR_TO_STACK register holding fp+0 against an
    iterator stored at fp-8 even though the runtime receives the actual fp+0
    pointer. An effectful iterator kfunc can then interpret memory outside
    the BPF stack as iterator state.
    
    Reject non-negative offsets before converting the offset to an spi. All
    valid stack objects begin at a negative offset from the frame pointer.
    
    Fixes: 06accc8779c1 ("bpf: add support for open-coded iterator loops")
    Signed-off-by: Xu Yunxiang <[email protected]>
    Signed-off-by: Andrii Nakryiko <[email protected]>
    Reviewed-by: Sun Jian <[email protected]>
    Link: https://lore.kernel.org/bpf/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Reject pkt arguments in mutating subprogs [+ + +]
Author: Emil Tsalapatis <[email protected]>
Date:   Tue Sep 22 17:20:22 2026 +0000

    bpf: Reject pkt arguments in mutating subprogs
    
    [ Upstream commit a6c1edfbe240e4377038a0e1d233ad81fde9a21b ]
    
    The verifier tracks changes in how PTR_TO_PACKET registers'
    bounds are modified across subprog boundaries. PTR_TO_PACKET
    registers are actually passed as PTR_TO_MEM, which is assumed
    valid for the entire call. This is not the case with packet memory,
    where a pskb_* call may invalidate its memory region.
    
    Reject BPF code that passes PTR_TO_PACKET pointers to subprogs that
    may mutate a packet. We cannot pass the pointer as a true PTR_TO_PACKET
    because we would also need to somehow pass the PTR_TO_PACKET_META
    or PTR_TO_PACKET_END to the subprog. Since we cannot avoid representing
    the pointer in the subprog as PTR_TO_MEM, only permit it if the
    subprog is guaranteed not to mutate the packet.
    
    Fixes: 80f281664f5a ("bpf: Support pointers in global func args")
    Reported-by: Nicholas Carlini <[email protected]>
    Suggested-by: Nicholas Carlini <[email protected]>
    Signed-off-by: Emil Tsalapatis <[email protected]>
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Restrict CO-RE poisoning to relocatable instructions [+ + +]
Author: Kumar Kartikeya Dwivedi <[email protected]>
Date:   Fri Sep 18 01:32:14 2026 +0200

    bpf: Restrict CO-RE poisoning to relocatable instructions
    
    [ Upstream commit 394ae398337c5f87e567f6cd63b937fc2b2f6ddc ]
    
    CO-RE relocation records can name any instruction offset. When a
    relocation cannot be resolved, bpf_core_patch_insn() currently poisons its
    target before checking whether that instruction is a valid relocation
    target. Malformed metadata can therefore replace jumps, calls, exits,
    register-source arithmetic, or non-immediate loads instead of failing at
    the relocation step.
    
    Handle poisoning only after the instruction has passed the same class and
    operand-form checks used for a resolved relocation. Route invalid forms
    through the existing diagnostic and return a hard error. Keep poisoning
    supported instructions, including both halves of a plain ldimm64, so an
    unresolved relocation in dead code remains valid.
    
    Extend bpf_core_poison_insn() to poison both halves of ldimm64, and return
    its status directly from each validated instruction case. This avoids
    routing the success path through a common label and leaves the helper free
    to report errors.
    
    The shared relocation code applies this restriction to both libbpf and
    in-kernel CO-RE.
    
    Fixes: d7a252708dbc ("libbpf: Improve handling of failed CO-RE relocations")
    Reported-by: Nicholas Carlini <[email protected]>
    Suggested-by: Nicholas Carlini <[email protected]>
    Signed-off-by: Kumar Kartikeya Dwivedi <[email protected]>
    Acked-by: Eduard Zingerman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Eduard Zingerman <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Skip unsettled links in link iterator [+ + +]
Author: Weiming Shi <[email protected]>
Date:   Tue Sep 15 01:02:07 2026 +0800

    bpf: Skip unsettled links in link iterator
    
    [ Upstream commit 50e80e2bb5e2be8515205b9c496b9640ddefa434 ]
    
    bpf_link_prime() inserts a link into link_idr before anon_inode_getfile()
    succeeds and before bpf_link_settle() publishes the ID in link->id.
    bpf_link_by_id() treats such an ID-zero link as unsettled, but the link
    iterator takes a reference without this check.
    
    If anon_inode_getfile() then fails, the creator removes the ID and frees
    its still-private link directly.  The iterator is left with a dangling
    reference and its next bpf_link_put() accesses freed memory.
    
    Treat ID-zero entries as transient in bpf_link_get_curr_or_next(), just as
    bpf_link_by_id() does.
    
      BUG: KASAN: slab-use-after-free in bpf_link_put
      Write of size 8 by task exp/384
      Call Trace:
      bpf_link_put                    kernel/bpf/syscall.c:3372
      bpf_link_seq_next               kernel/bpf/link_iter.c:33
      bpf_seq_read                    kernel/bpf/bpf_iter.c:158
      vfs_read                        fs/read_write.c:572
      ksys_read                       fs/read_write.c:716
      do_syscall_64                   arch/x86/entry/syscall_64.c:84
      entry_SYSCALL_64_after_hwframe  arch/x86/entry/entry_64.S:121
      Kernel panic - not syncing: KASAN: panic_on_warn set ...
    
    Fixes: 9f8836127308 ("bpf: Add bpf_link iterator")
    Reported-by: Xiang Mei <[email protected]>
    Signed-off-by: Weiming Shi <[email protected]>
    Signed-off-by: Andrii Nakryiko <[email protected]>
    Link: https://lore.kernel.org/bpf/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element [+ + +]
Author: Donggeun Yoo <[email protected]>
Date:   Thu Sep 24 19:23:20 2026 +0900

    bpf: Zero-fill other CPUs when BPF_F_CPU creates a per-cpu hash element
    
    [ Upstream commit c3a66e5f5bab3912e9f84223c5a982bf4333d5a1 ]
    
    pcpu_init_value() initializes the per-cpu area of a newly created
    [lru_]percpu_hash element.  The area is recycled, so when the value
    comes from a BPF program (onallcpus == false) it writes the running
    CPU's slot and zeroes the rest.
    
    bpf_percpu_hash_update() passes onallcpus == true, which delegates to
    pcpu_copy_value().  pcpu_copy_value() writes only the CPU named in
    map_flags when BPF_F_CPU is set, so on the create path the other slots
    keep the recycled element's values:
    
      update(k1, 0xdeadc0de, BPF_F_ALL_CPUS)  every CPU holds 0xdeadc0de
      delete(k1)                              element back on the freelist
      update(k2, 0xc0ffee, BPF_F_CPU | 0)     creates, writes CPU 0 only
      lookup(k2)                              CPU 0 0xc0ffee, rest 0xdeadc0de
    
    Zero-fill the other CPUs on that arm too.
    
    Fixes: c6936161fd55 ("bpf: Add BPF_F_CPU and BPF_F_ALL_CPUS flags support for percpu_hash and lru_percpu_hash maps")
    Signed-off-by: Donggeun Yoo <[email protected]>
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

 
bridge: check llc_mac_hdr_init() return value in br_send_bpdu() [+ + +]
Author: Eric Dumazet <[email protected]>
Date:   Thu Sep 24 08:29:49 2026 +0000

    bridge: check llc_mac_hdr_init() return value in br_send_bpdu()
    
    [ Upstream commit ac704ff08e511c87643799c385f55ecd69b85e03 ]
    
    If llc_mac_hdr_init() fails (for instance if the port device type does
    not support LLC or dev_hard_header() fails), br_send_bpdu() should drop
    the skb instead of resetting the mac header to the LLC payload and
    transmitting a malformed frame.
    
    Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
    Closes: https://lore.kernel.org/netdev/[email protected]/
    Cc: Nikolay Aleksandrov <[email protected]>
    Cc: Ido Schimmel <[email protected]>
    Cc: [email protected]
    Signed-off-by: Eric Dumazet <[email protected]>
    Acked-by: Nikolay Aleksandrov <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
cgroup/cpuset: Return PERR_NOCPUS in remote_partition_enable() on subpartitions_cpus conflict [+ + +]
Author: Hui Peng <[email protected]>
Date:   Thu Sep 24 04:27:27 2026 +0000

    cgroup/cpuset: Return PERR_NOCPUS in remote_partition_enable() on subpartitions_cpus conflict
    
    commit 31c88350b7dd1522792f726f79607f31bb55c50f upstream.
    
    When a remote partition is created underneath an existing local partition
    via a non-partition (PRS_MEMBER) intermediate cgroup, update_prstate() sees
    parent->partition_root_state == PRS_MEMBER and calls
    remote_partition_enable().
    
    Commit 86888c7bd117 ("cgroup/cpuset: Add warnings to catch inconsistency
    in exclusive CPUs") replaced the cpumask_intersects(tmp->new_cpus,
    subpartitions_cpus) error check in remote_partition_enable() with
    WARN_ON_ONCE(). As a result, remote_partition_enable() emits a warning
    and proceeds to enable the remote partition on CPUs that are already
    owned by the ancestor local partition in subpartitions_cpus.
    
    This can be reproduced on Linux 7.3.0-rc3 with:
    
      mkdir -p /tmp/cg1
      mount -t cgroup2 none /tmp/cg1
      echo "+cpuset" > /tmp/cg1/cgroup.subtree_control
    
      mkdir /tmp/cg1/A
      echo 1 > /tmp/cg1/A/cpuset.cpus
      echo 1 > /tmp/cg1/A/cpuset.cpus.exclusive
      echo root > /tmp/cg1/A/cpuset.cpus.partition
      echo "+cpuset" > /tmp/cg1/A/cgroup.subtree_control
    
      mkdir /tmp/cg1/A/B
      echo 1 > /tmp/cg1/A/B/cpuset.cpus
      echo 1 > /tmp/cg1/A/B/cpuset.cpus.exclusive
      echo "+cpuset" > /tmp/cg1/A/B/cgroup.subtree_control
    
      mkdir /tmp/cg1/A/B/D
      echo 1 > /tmp/cg1/A/B/D/cpuset.cpus
      echo 1 > /tmp/cg1/A/B/D/cpuset.cpus.exclusive
      echo root > /tmp/cg1/A/B/D/cpuset.cpus.partition
    
    which triggers:
    
      WARNING: kernel/cgroup/cpuset.c:1594 at remote_partition_enable+0x1c1/0x300
    
    and leaves both /tmp/cg1/A and /tmp/cg1/A/B/D as active root partitions
    claiming exclusive CPU 1.
    
    Fix this by returning PERR_NOCPUS when tmp->new_cpus intersects
    subpartitions_cpus in remote_partition_enable(), matching the error code
    used by remote_cpus_update() for the same subpartitions_cpus conflict, and
    add a regression test case to
    tools/testing/selftests/cgroup/test_cpuset_prs.sh.
    
    Tested in QEMU on Linux 7.3.0-rc3 using the reproducer above and
    tools/testing/selftests/cgroup/test_cpuset_prs.sh.
    
    Fixes: 86888c7bd117 ("cgroup/cpuset: Add warnings to catch inconsistency in exclusive CPUs")
    Suggested-by: Guopeng Zhang <[email protected]>
    Cc: [email protected]
    Assisted-by: LLM
    Signed-off-by: Hui Peng <[email protected]>
    Reviewed-by: Waiman Long <[email protected]>
    Signed-off-by: Tejun Heo <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
cgroup/pids: Restore pids.events notifications in local mode [+ + +]
Author: Guopeng Zhang <[email protected]>
Date:   Thu Sep 24 17:15:36 2026 +0800

    cgroup/pids: Restore pids.events notifications in local mode
    
    commit 1765a153d985c231357145e26798f9408db10e42 upstream.
    
    A fork rejected by the pids controller increments the counter reported by
    pids.events. When local event accounting is selected, however, pids_event()
    returns after notifying only events_local_file, leaving pids.events pollers
    asleep.
    
    On legacy hierarchies, pids.events.local does not exist. With
    pids_localevents, pids.events reports the same local counter. In both
    cases, pids.events changes without generating a notification.
    
    This can be reproduced with a pids_localevents mount:
    
        mkdir /tmp/test
        mount -t cgroup2 -o pids_localevents none /tmp/test
        mkdir /tmp/test/t
        echo 1 > /tmp/test/t/pids.max
        cat /tmp/test/t/pids.events                 # max 0
        timeout 3 inotifywait -e modify /tmp/test/t/pids.events &
        sh -c 'echo $$ > /tmp/test/t/cgroup.procs; (true &)' 2>/dev/null
        wait
        cat /tmp/test/t/pids.events                 # max 1
    
    Without this patch, inotifywait times out without reporting an event.
    Notify pids.events before returning from the local event path.
    
    Fixes: 3f26a885a068 ("cgroup/pids: Add pids.events.local")
    Cc: [email protected] # v6.11+
    Signed-off-by: Guopeng Zhang <[email protected]>
    Signed-off-by: Tejun Heo <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
crypto: s390/hmac - Generate intermediate CV for API partial block handling [+ + +]
Author: Holger Dengler <[email protected]>
Date:   Fri Aug 14 16:20:10 2026 +0200

    crypto: s390/hmac - Generate intermediate CV for API partial block handling
    
    commit 10396a2d6d41d594975b6ece712278570c3c970c upstream.
    
    The API partial block handling requires a intermediate chaining
    value (CV). The internal function hash_data() sets the function code
    correctly, so also call cpacf_kimd() instruction for intermediate CV
    generation, as cpacf_klmd() always generate the final hash value.
    
    Cc: [email protected] # 6.15+
    Fixes: 08811169ac01 ("crypto: s390/hmac - Use API partial block handling")
    Signed-off-by: Holger Dengler <[email protected]>
    Reviewed-by: Harald Freudenberger <[email protected]>
    Signed-off-by: Herbert Xu <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
dpll: use exact lookup for reference sync pin id [+ + +]
Author: Ivan Vecera <[email protected]>
Date:   Thu Sep 17 16:37:36 2026 +0200

    dpll: use exact lookup for reference sync pin id
    
    [ Upstream commit 7cce782d8327b7291334c4a304cf3fd909a74d9d ]
    
    dpll_pin_ref_sync_state_set() looks up the reference sync pin in the
    pin->ref_sync_pins xarray, which is keyed by the sync pin's id (see
    dpll_pin_ref_sync_pair_add() using xa_insert() with ref_sync_pin->id).
    The pin id to operate on is supplied by userspace via DPLL_A_PIN_ID.
    
    The lookup however used xa_find() with a ULONG_MAX limit, which returns
    the first present entry with an index greater than or equal to the
    requested id, not the entry stored exactly at that id. If userspace
    passes an id that is not paired as a reference sync pin, but another
    pin with a higher id is present in the xarray, xa_find() silently
    returns that wrong pin and the subsequent ref_sync_set() operates on
    it. The request only fails when the given id is larger than every
    present key.
    
    Use xa_load() for an exact-key lookup instead, mirroring the deletion
    path in dpll_pin_ref_sync_pair_del().
    
    Fixes: 58256a26bfb3 ("dpll: add reference sync get/set")
    Signed-off-by: Ivan Vecera <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
drm/amd/display: Atomize IRQ register read/modify/write ops [+ + +]
Author: Leo Li <[email protected]>
Date:   Wed Sep 23 11:43:30 2026 -0500

    drm/amd/display: Atomize IRQ register read/modify/write ops
    
    [ Upstream commit 63e19ef3ddab806c472748c825f4dc88dcd994e8 ]
    
    [Why]
    
    The OTG_GLOBAL_SYNC_STATUS register controls various HW IRQ sources for
    the output timing generator (OTG). VUPDATE_NO_LOCK is one of them.
    
    To enable the IRQ, driver sets the VUPDATE_NO_LOCK_EN bit in the
    GLOBAL_SYNC_STATUS register.
    
    To ack the IRQ after it fires, the driver sets the VUPDATE_NO_LOCK_CLEAR
    bit in the same GLOBAL_SYNC_STATUS register.
    
    The bit sets are done through read/modify/write operations, which are
    not atomic. Thus, the following race is possible:
    
        Thread A:                       IRQ handler:
                                        *HW IRQ fires*
        # IRQ disable
        val = read(GLOBAL_SYNC_STATUS)
        unset(val, VUPDATE_NO_LOCK_EN)
        write(val, GLOBAL_SYNC_STATUS)
                                        # ACK reads VUPDATE_NO_LOCK_EN unset
                                        val1 = read(GLOBAL_SYNC_STATUS)
                                        set(val1, VUPDATE_NO_LOCK_CLEAR)
        # IRQ enable
        val = read(GLOBAL_SYNC_STATUS)
        set(val, VUPDATE_NO_LOCK_EN)
        write(val, GLOBAL_SYNC_STATUS)
                                        # BAD! clears VUPDATE_NO_LOCK_EN
                                        write(val1, GLOBAL_SYNC_STATUS)
    
    Regarding the tagged Fixes: change, it appears the change made this race
    more likely to occur. Since VUPDATE_NO_LOCK is now the sole IRQ source
    for vblank handling, a single race on high refresh panels can lead to a
    time out.
    
    [How]
    
    The GLOBAL_SYNC_STATUS register is only one example, other IRQ control
    registers also share the same scheme. On top of GLOBAL_SYNC_STATUS,
    let's clean up those as well.
    
    To keep things simple, Let's atomize the IRQ rmw ops via a single
    driver-wide spinlock. Due to the small scope of this lock, it is
    unlikely to cause noticeable overhead on top of all the existing locking
    within the IRQ set/handle paths.
    
    Since DM is responsible for locking, wrap dc_interrupt_set/ack with the
    spinlock in the new amdgpu_dm_irq_set/ack functions. Migrate/drop all
    references in DM to dc_interrupt_set/ack to use amdgpu_dm_irq_set/ack
    instead.
    
    Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5616
    Fixes: c87e6635d2db ("drm/amd/display: consolidate DCN vblank/flip handling onto vupdate_no_lock")
    Reviewed-by: Mario Limonciello <[email protected]>
    Signed-off-by: Leo Li <[email protected]>
    Signed-off-by: Chenyu Chen <[email protected]>
    Tested-by: Daniel Wheeler <[email protected]>
    Signed-off-by: Alex Deucher <[email protected]>
    (cherry picked from commit 70de0a0216583a53c946155f8c8adedfdca6b4e7)
    Cc: [email protected]
    (cherry picked from commit 63e19ef3ddab806c472748c825f4dc88dcd994e8)
    Modified for unit tests not present in 7.2.y
    Signed-off-by: Mario Limonciello <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>
drm/amd/display: Bump frame warning limit for clang builds of dml [+ + +]
Author: Ivan Lipski <[email protected]>
Date:   Fri Aug 21 00:04:53 2026 -0400

    drm/amd/display: Bump frame warning limit for clang builds of dml
    
    commit 18779dd84515db093fedb4ebaf0998c9b165a5fb upstream.
    
    [Why&How]
    When building the DML files with clang without any sanitizer or LTO,
    the following -Wframe-larger-than errors break the build under
    CONFIG_WERROR:
    
      display_mode_vba_30.c: error: stack frame size (2512) exceeds limit
        (2048) in 'dml30_ModeSupportAndSystemConfigurationFull'
      display_mode_vba_31.c: error: stack frame size (2416) exceeds limit
        (2048) in 'dml31_ModeSupportAndSystemConfigurationFull'
      display_mode_vba_314.c: error: stack frame size (2392) exceeds limit
        (2048) in 'dml314_ModeSupportAndSystemConfigurationFull'
    
    Clang consistently spills more than gcc, pushing the frame past the 2048
    byte limit.
    
    Apply an existing approach of increasing the warn stack size to the
    non-sanitizer path so plain clang builds use a 3072 byte limit.
    
    Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5642
    Signed-off-by: Ivan Lipski <[email protected]>
    Signed-off-by: Alex Deucher <[email protected]>
    (cherry picked from commit 21711b6e66bb7b41b1aec67b2d99aafe768c8fcb)
    Cc: [email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/amd/display: Fix dc stream excess put in dm_update_crtc_state() [+ + +]
Author: Wentao Liang <[email protected]>
Date:   Wed Sep 16 09:49:56 2026 +0000

    drm/amd/display: Fix dc stream excess put in dm_update_crtc_state()
    
    commit c5fd4eaad50d620c7e09ac2082b2fb55ee54170e upstream.
    
    In dm_update_crtc_state(), when a modeset is required the newly created
    stream is stored in dm_new_crtc_state->stream and an extra reference is
    taken with dc_stream_retain().  The reference returned by
    create_validate_stream_for_sink() is released as an extra reference at
    the skip_modeset label, leaving the stream owned by the new CRTC state.
    
    If amdgpu_dm_check_crtc_color_mgmt() fails afterwards, the code jumps
    to the fail label which releases new_stream again.  Since the extra
    reference was already released at skip_modeset, this drops the
    reference owned by dm_new_crtc_state->stream and the stream is
    released while the atomic state still points to it, leading to a
    premature free of the dc stream.
    
    Set new_stream to NULL after releasing the extra reference at the
    skip_modeset label so that a later goto fail cannot release the
    reference owned by the new CRTC state.
    
    Fixes: 7cd4b70091a5 ("drm/amd/display: Rework CRTC color management")
    Signed-off-by: Wentao Liang <[email protected]>
    Signed-off-by: Alex Deucher <[email protected]>
    (cherry picked from commit 102a47065a62dc8f6bbbb47cf082a2934282eb08)
    Cc: [email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/amd/display: Relax DML frame limit with UBSAN [+ + +]
Author: Alex Hung <[email protected]>
Date:   Tue Sep 29 15:30:14 2026 -0400

    drm/amd/display: Relax DML frame limit with UBSAN
    
    [ Upstream commit 0fd5e9ddf362b1253b3c94b858371f90400e5c4a ]
    
    [WHY]
    UBSAN instrumentation adds checks and handler calls and increases
    stack usage in the large DML calculation functions, similar to
    KASAN and KCSAN. With UBSAN enabled these files exceed the default
    -Wframe-larger-than limit and fail to build when -Werror is in effect.
    
    Reproduced with LLVM (make LLVM=1, clang 19.1.1), CONFIG_UBSAN=y,
    CONFIG_GCOV_PROFILE_ALL=y and CONFIG_DRM_AMDGPU_WERROR=y on x86_64.
    
    [HOW]
    Include CONFIG_UBSAN in the sanitizer check that selects the
    higher per-file frame warning limit in the dml and dml2_0
    Makefiles.
    
    Suggested-by: Leo Li <[email protected]>
    Assisted-by: Copilot:Claude-Opus-5.5
    Signed-off-by: Alex Hung <[email protected]>
    Signed-off-by: Alex Deucher <[email protected]>
    (cherry picked from commit ebf8b0fd8508b744f85a8eee82b745b1d3502dd0)
    Cc: [email protected]
    [ Changed sanitizer checks to test for a nonempty filter result over space-separated configuration values. ]
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/amd/display: Remove sink usage from DPMS [+ + +]
Author: Dominik Kaszewski <[email protected]>
Date:   Fri Jun 19 11:22:07 2026 +0200

    drm/amd/display: Remove sink usage from DPMS
    
    [ Upstream commit 8fa813b7fd1eccbed13e126166cf764f1f11a7d3 ]
    
    [Why]
    stream->sink is optional and can be null, so should always be checked
    before dereference. Additionally, most of its usage in DPMS sequences
    is for stream->sink->link, which can be replaced with stream->link,
    as the two should always be the same.
    
    [How]
    * Replace stream->sink->link in DPMS on/off
    * Add assert to USB4 BW allocation where sink is required
    * Avoid inconsistencies in resource access, e.g. don't repeat
    stream->link after it was already saved to a local variable
    * Pull out effective VPG calculation to helper getter
    * Formatting fixes
    
    Reviewed-by: Nicholas Kazlauskas <[email protected]>
    Signed-off-by: Dominik Kaszewski <[email protected]>
    Signed-off-by: George Zhang <[email protected]>
    Signed-off-by: Alex Deucher <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
drm/amdgpu/userq: fix double jiffies conversion in hang detect timeout [+ + +]
Author: Sunil Khatri <[email protected]>
Date:   Thu Sep 17 18:56:45 2026 +0530

    drm/amdgpu/userq: fix double jiffies conversion in hang detect timeout
    
    commit cd195f1616b2bb5fb7765465326c4d5d64a620a0 upstream.
    
    Function amdgpu_userq_start_hang_detect_work() calls msecs_to_jiffies()
    on adev->gfx_timeout/compute_timeout/sdma_timeout before arming
    hang_detect_work. These timeout values already hold jiffies values from
    amdgpu_device_get_job_timeout_settings() at device init.
    
    This silently shrinks the real hang-detect deadline to (2 * HZ) ms
    instead of the intended timeout. e.g. 500ms instead of the 2000ms
    default on a CONFIG_HZ=250 kernel, only coincidentally correct at
    HZ=1000. The shortened window is easily exceeded by ordinary
    fence-completion latency, causing hang_detect_work to fire and
    trigger a per-queue or full GPU reset for queues that are not
    actually hung.
    
    Pass the jiffies value directly to queue_delayed_work() instead of
    converting it a second time.
    
    Fixes: fc3336be9c62 ("drm/amd/amdgpu: Add independent hang detect work for user queue fence")
    Signed-off-by: Sunil Khatri <[email protected]>
    Reviewed-by: Christian König <[email protected]>
    Signed-off-by: Alex Deucher <[email protected]>
    (cherry picked from commit 13d44ca033cb74756c2aef0ade54a75cdf2f6271)
    Cc: [email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
drm/amdgpu/vcn4.0.3: fix video_timeout unit mismatch in jpeg reset wait [+ + +]
Author: Sunil Khatri <[email protected]>
Date:   Thu Sep 17 18:56:45 2026 +0530

    drm/amdgpu/vcn4.0.3: fix video_timeout unit mismatch in jpeg reset wait
    
    commit 6b13ddbf5bb8deec337f9b4887f579097a953f6a upstream.
    
    vcn_v4_0_3_reset_jpeg_pre_helper() passes adev->video_timeout directly
    to amdgpu_fence_wait_polling(), whose timeout parameter is documented
    and implemented in usecs (busy-wait loop decrementing by udelay(2)).
    
    adev->video_timeout is set in jiffies by
    amdgpu_device_get_job_timeout_settings(), via msecs_to_jiffies().
    Passing it unconverted means the intended ~2s wait for outstanding
    JPEG fences to complete before the JPEG queue is torn down actually
    lasts only a couple of microseconds (HZ jiffies interpreted as usecs),
    so pending jobs are almost never given a real chance to finish before
    the reset path forces completion in the following helper.
    
    Convert the jiffies value to usecs with jiffies_to_usecs() before
    passing it to amdgpu_fence_wait_polling().
    
    Fixes: d25c67fd9d6f ("drm/amdgpu/vcn4.0.3: rework reset handling")
    Cc: Jesse.Zhang <[email protected]>
    Assisted-by: Claude:claude-sonnet-5
    Signed-off-by: Sunil Khatri <[email protected]>
    Reviewed-by: Christian König <[email protected]>
    Signed-off-by: Alex Deucher <[email protected]>
    (cherry picked from commit 5feabbd673c10ebee22b880e4d812f08974d2ef7)
    Cc: [email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
drm/amdgpu/vcn5.0.1: fix video_timeout unit mismatch in jpeg reset wait [+ + +]
Author: Sunil Khatri <[email protected]>
Date:   Thu Sep 17 18:56:45 2026 +0530

    drm/amdgpu/vcn5.0.1: fix video_timeout unit mismatch in jpeg reset wait
    
    commit f952ed353a27b46c86c9525a39b7a642850b8139 upstream.
    
    vcn_v5_0_1_reset_jpeg_pre_helper() passes adev->video_timeout directly
    to amdgpu_fence_wait_polling(), whose timeout parameter is documented
    and implemented in usecs (busy-wait loop decrementing by udelay(2)).
    
    adev->video_timeout is set in jiffies by
    amdgpu_device_get_job_timeout_settings(), via msecs_to_jiffies().
    Passing it unconverted means the intended ~2s wait for outstanding
    JPEG fences to complete before the JPEG queue is torn down actually
    lasts only a couple of microseconds (HZ jiffies interpreted as usecs),
    so pending jobs are almost never given a real chance to finish before
    the reset path forces completion in the following helper.
    
    Convert the jiffies value to usecs with jiffies_to_usecs() before
    passing it to amdgpu_fence_wait_polling().
    
    Fixes: fab47d2db5ca ("drm/amdgpu/vcn5.0.1: rework reset handling")
    Cc: Jesse.Zhang <[email protected]>
    Assisted-by: Claude:claude-sonnet-5
    Signed-off-by: Sunil Khatri <[email protected]>
    Reviewed-by: Christian König <[email protected]>
    Signed-off-by: Alex Deucher <[email protected]>
    (cherry picked from commit b8334fec8b90ebffcaa01001a23edca9f29a05e9)
    Cc: [email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
drm/amdgpu: Fix acpi device leak in amdgpu_acpi_enumerate_xcc() [+ + +]
Author: Wentao Liang <[email protected]>
Date:   Wed Sep 16 09:54:32 2026 +0000

    drm/amdgpu: Fix acpi device leak in amdgpu_acpi_enumerate_xcc()
    
    commit a997baa61179b450bd55c4810c7ccfed3b753a54 upstream.
    
    amdgpu_acpi_enumerate_xcc() looks up each XCC ACPI device with
    acpi_dev_get_first_match_dev(), which takes a reference to the device.
    The reference is dropped with acpi_dev_put() after the XCC info is
    initialized, but if the kzalloc_obj() allocation of the XCC info fails
    the function returns -ENOMEM without releasing the reference, leaking
    the last reference to the ACPI device.
    
    Drop the ACPI device reference on the allocation failure path before
    returning.
    
    Fixes: 4d5275ab0b18 ("drm/amdgpu: Add parsing of acpi xcc objects")
    Reviewed-by: Lijo Lazar <[email protected]>
    Signed-off-by: Wentao Liang <[email protected]>
    Signed-off-by: Alex Deucher <[email protected]>
    (cherry picked from commit 9211ef48b31ec66999cf55e04d0cbc60cd855fd5)
    Cc: [email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/amdgpu: Fix last_update fence leak in amdgpu_vm_init() [+ + +]
Author: Wentao Liang <[email protected]>
Date:   Wed Sep 16 09:55:36 2026 +0000

    drm/amdgpu: Fix last_update fence leak in amdgpu_vm_init()
    
    commit b4f7b4459b1b155e4c4977a6482b5df2cf08758c upstream.
    
    amdgpu_vm_init() initializes vm->last_update, vm->last_unlocked and
    vm->last_tlb_flush with references to the stub fence taken via
    dma_fence_get_stub().  The error label at the end of the function
    releases the last_unlocked and last_tlb_flush references with
    dma_fence_put(), but the reference stored in vm->last_update is never
    dropped, so whenever the page table root creation, the reservation of
    the root BO or the PASID registration fails, the stub fence reference
    leaks.
    
    Drop the vm->last_update reference together with the other stub fence
    references on the error path.
    
    Fixes: 187916e6ed9d ("drm/amdgpu: install stub fence into potential unused fence pointers")
    Signed-off-by: Wentao Liang <[email protected]>
    Signed-off-by: Alex Deucher <[email protected]>
    (cherry picked from commit e7979c84fc05a176bdf855ee664871b1648404c9)
    Cc: [email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/amdgpu: Fix runtime PM leak in amdgpu_debugfs_test_ib_show() [+ + +]
Author: Wentao Liang <[email protected]>
Date:   Wed Sep 16 09:58:05 2026 +0000

    drm/amdgpu: Fix runtime PM leak in amdgpu_debugfs_test_ib_show()
    
    commit 2b86ab1bd6673c525adda88819d7658ba9e784ec upstream.
    
    amdgpu_debugfs_test_ib_show() resumes the device with
    pm_runtime_get_sync() before taking the reset domain semaphore with
    down_write_killable().  If the write lock acquisition is interrupted,
    the function returns without calling pm_runtime_put_autosuspend(),
    leaking the runtime PM reference acquired for the device and keeping
    the GPU awake.
    
    Drop the runtime PM reference on the interrupted down_write_killable()
    error path before returning.
    
    Fixes: 6049db43d6dd ("drm/amdgpu: change reset lock from mutex to rw_semaphore")
    Signed-off-by: Wentao Liang <[email protected]>
    Signed-off-by: Alex Deucher <[email protected]>
    (cherry picked from commit ec30a576c2d4c0364549e6c04218f50704ef56c8)
    Cc: [email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/amdgpu: Fix vmid_wait fence leak in amdgpu_ring_init() [+ + +]
Author: Wentao Liang <[email protected]>
Date:   Wed Sep 16 10:01:39 2026 +0000

    drm/amdgpu: Fix vmid_wait fence leak in amdgpu_ring_init()
    
    commit aea841bc62a76242396610d22d8ff40c13065f64 upstream.
    
    amdgpu_ring_init() initializes ring->vmid_wait with a reference to the
    stub fence taken via dma_fence_get_stub().  When a later step of the
    initialization fails, e.g. amdgpu_fence_driver_init_ring(), a writeback
    slot allocation or the ring buffer allocation, the function returns an
    error without releasing the stub fence reference and the reference is
    leaked if the ring is torn down without amdgpu_ring_fini().
    
    Move the stub fence assignment to the end of the initialization, right
    before the ring is registered with the GPU scheduler, where no further
    failure is possible.  The stub fence is only consumed by command
    submission handling in amdgpu_ids.c once the ring is up and running, so
    nothing reads it during the error-prone part of the initialization.
    
    Fixes: 48e9fbd1a284 ("drm/amdgpu: initialize the vmid_wait with the stub fence")
    Signed-off-by: Wentao Liang <[email protected]>
    Signed-off-by: Alex Deucher <[email protected]>
    (cherry picked from commit f2b96986851203e9c50ca0d13aaa3581ca3e8ebd)
    Cc: [email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/amdgpu: move userq fence wait out of signalling section [+ + +]
Author: Prike Liang <[email protected]>
Date:   Fri Jul 31 11:44:37 2026 +0800

    drm/amdgpu: move userq fence wait out of signalling section
    
    commit 3022bdfe3e6d776e9273d6892f7c193138ca0666 upstream.
    
    The eviction fence suspend worker waits for every pending userq fence
    from inside a dma_fence_begin_signalling() critical section. Waiting on
    another DMA fence while responsible for signalling one violates the
    cross-driver fence contract and is reported by lockdep as a
    dma_fence_map dependency.
    
    Move the wait before dma_fence_begin_signalling(). Keep userq_mutex held
    so queue lifetime remains stable while inspecting last_fence.
    
    Fixes: fc61df151617 ("drm/amdgpu: annotate eviction fence signaling path")
    Signed-off-by: Prike Liang <[email protected]>
    Reviewed-by: Vitaly Prosyak <[email protected]>
    Signed-off-by: Alex Deucher <[email protected]>
    (cherry picked from commit 3bd4fbc5ed89621340b5cd249869092691a9c81f)
    Cc: [email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
drm/amdkfd: fix use-after-free and multi-container gap in kfd_dev_mapping [+ + +]
Author: Asad Kamal <[email protected]>
Date:   Fri Aug 28 15:59:21 2026 +0800

    drm/amdkfd: fix use-after-free and multi-container gap in kfd_dev_mapping
    
    commit c3a31087b1c8df679b653c1a09d9abe5ff7ec8ef upstream.
    
    kfd_dev_mapping caches the address_space of the first /dev/kfd opener
    so that the GPU reset path can call unmap_mapping_range() to zap all
    userspace mappings of doorbell and MMIO ranges.  This design has two
    bugs that both manifest under SRIOV with multiple containers:
    
    1. Use-after-free / rwsem deadlock.  The cached pointer refers to an
       inode owned by the first opener's container.  When that container
       exits and its inode is released, kfd_dev_mapping becomes a dangling
       pointer.  A subsequent GPU reset dereferences it inside
       unmap_mapping_range(), which takes i_mmap_rwsem on the freed inode,
       causing a hard hang observable as an uninterruptible rwsem wait.
    
    2. Multi-container gap.  Only the first opener's address_space is cached;
       VMAs created by later openers live in a different address_space and
       are never reached by unmap_mapping_range().  After a GPU reset those
       stale mappings keep doorbell and MMIO pages accessible to guest
       userspace with no GPU behind them, risking PCIe transaction timeouts
       and NMI panics.
    
    Fix both bugs with the same approach used by DRM core (drm_drv.c):
    create a private pseudo-filesystem at module init time and allocate one
    anonymous inode from it.  In kfd_open() redirect every opener's
    file->f_mapping to that inode's address_space.  The inode is
    module-owned, lives exactly as long as the amdgpu module, and collects
    VMAs from all openers in one address_space.  A single
    unmap_mapping_range() call in the reset path then correctly reaches
    every container's mappings with no dangling pointer risk.
    
    The hang manifests as an NMI backtrace on the GPU reset workqueue stuck
    spinning in rwsem_down_read_slowpath() with a corrupted i_mmap_rwsem:
    
      Workqueue: amdgpu-reset-dev xgpu_ai_mailbox_flr_work [amdgpu]
      Call Trace:
       <TASK>
       kvm_wait+0x1f/0x40
       __pv_queued_spin_lock_slowpath+0x31d/0x3a0
       _raw_spin_lock_irq+0x51/0x80
       rwsem_down_read_slowpath+0xb3/0x550
       down_read+0x48/0xd0
       unmap_mapping_range+0x71/0x140
       kfd_dev_unmap_mapping_range+0x5b/0x140 [amdgpu]
       amdgpu_amdkfd_clear_kfd_mapping+0xd8/0x190 [amdgpu]
       amdgpu_device_gpu_recover+0x232/0x450 [amdgpu]
       xgpu_ai_mailbox_flr_work+0xb5/0xc0 [amdgpu]
       process_one_work+0x18e/0x3e0
       worker_thread+0x2e3/0x420
       kthread+0x10a/0x230
    
    Fixes: 70cadefcc616 ("drm/amdgpu: unmap all user mappings of framebuffer and doorbell before mode1 reset")
    Signed-off-by: Asad Kamal <[email protected]>
    Reviewed-by: Lijo Lazar <[email protected]>
    Signed-off-by: Alex Deucher <[email protected]>
    (cherry picked from commit 1128b4a52de1572e87431de837fd9850cb99542c)
    Cc: [email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
drm/bridge: samsung-dsim: fix TE GPIO lifetime for host attach [+ + +]
Author: Li Youhong <[email protected]>
Date:   Fri Sep 4 09:49:58 2026 +0800

    drm/bridge: samsung-dsim: fix TE GPIO lifetime for host attach
    
    [ Upstream commit ada667890773e033d2f40dc94176e3beb930b516 ]
    
    When the Exynos DSI driver was generalized into samsung-dsim, the TE
    GPIO acquisition was switched from gpiod_get_optional() to
    devm_gpiod_get_optional() while keeping the matching gpiod_put() calls.
    That combination is wrong for a managed descriptor.
    
    However, dropping the puts and keeping the managed get is also wrong:
    samsung_dsim_register_te_irq() runs from the DSI host attach callback,
    and host detach/reattach can happen without destroying the device that
    owns the managed action. A second attach would then request the GPIO
    again without having released it.
    
    Switch back to a non-managed gpiod_get_optional() and keep the explicit
    gpiod_put() on the request_irq() error path and in
    samsung_dsim_unregister_te_irq().
    
    Fixes: e7447128ca4a ("drm: bridge: Generalize Exynos-DSI driver into a Samsung DSIM bridge")
    Suggested-by: Luca Ceresoli <[email protected]>
    Signed-off-by: Li Youhong <[email protected]>
    Reviewed-by: Luca Ceresoli <[email protected]>
    Tested-by: Luca Ceresoli <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    [Luca: remove unnecessary comment]
    Signed-off-by: Luca Ceresoli <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
drm/client: fix restore of partially initialized client [+ + +]
Author: shechenglong <[email protected]>
Date:   Mon Sep 7 11:51:47 2026 +0800

    drm/client: fix restore of partially initialized client
    
    commit 1fca688e9443003e33cf30453e7a7560367656c9 upstream.
    
    I got a null-ptr-deref report when closing a DRM file descriptor:
    
    WARNING: drivers/gpu/drm/drm_atomic.c:2031 at
    __drm_atomic_helper_set_config+0x18e/0x1b0 [drm]
    
    Call Trace:
    drm_client_modeset_commit_atomic+0x16b/0x220 [drm]
    drm_client_modeset_commit_locked+0x56/0x160 [drm]
    drm_client_modeset_commit+0x21/0x40 [drm]
    __drm_fb_helper_restore_fbdev_mode_unlocked.part.0+0x7b/0x80
    drm_fbdev_client_restore+0xe/0x20 [drm_client_lib]
    drm_client_dev_restore+0x9f/0xc0 [drm]
    drm_release+0xc5/0xe0 [drm]
    
    The warning is followed by a NULL pointer dereference:
    
    BUG: kernel NULL pointer dereference, address: 0000000000000008
    
    RIP:
    __drm_fb_helper_restore_fbdev_mode_unlocked.part.0+0x41/0x80
    [drm_kms_helper]
    
    Call Trace:
    drm_fbdev_client_restore+0xe/0x20 [drm_client_lib]
    drm_client_dev_restore+0x9f/0xc0 [drm]
    drm_release+0xc5/0xe0 [drm]
    __fput+0xdc/0x2b0
    __x64_sys_close+0x39/0x80
    do_syscall_64+0x8d/0x460
    entry_SYSCALL_64_after_hwframe+0x76/0x7e
    
    drm_client_register() adds the DRM client to the device client list
    before invoking the initial hotplug callback. If the hotplug callback
    fails, the client remains registered.
    
    For the fbdev client, a failure during drm_fb_helper_initial_config()
    causes the partially initialized fbdev helper to be cleaned up.
    drm_fb_helper_fini() releases fb_helper->info and leaves it NULL.
    
    The fbdev client therefore remains registered even though there is no
    fully initialized framebuffer device.
    
    Later, when userspace closes the DRM file descriptor, drm_release()
    can invoke the restore callbacks of registered DRM clients:
    
    drm_release()
    drm_client_dev_restore()
    drm_fbdev_client_restore()
    drm_fb_helper_restore_fbdev_mode_unlocked()
    
    drm_fbdev_client_restore() currently restores the fbdev state
    unconditionally. For a partially initialized fbdev client this can
    submit an incomplete modeset state and subsequently access fbdev
    state which has not been initialized, resulting in the warning and
    NULL pointer dereference above.
    
    drm_fbdev_client_unregister() already uses fb_helper->info to
    distinguish a fully probed framebuffer device from a partially
    initialized client.
    
    Use the same condition in drm_fbdev_client_restore() and skip restore
    if no framebuffer device has been successfully initialized.
    
    Signed-off-by: shechenglong <[email protected]>
    Reviewed-by: Thomas Zimmermann <[email protected]>
    Fixes: 5d08c44e47b9 ("drm/fbdev: Add memory-agnostic fbdev client")
    Signed-off-by: Thomas Zimmermann <[email protected]>
    Cc: <[email protected]> # v6.13+
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
drm/i915/dp_mst: Fix configuring FEC for a disconnected stream [+ + +]
Author: Imre Deak <[email protected]>
Date:   Mon Sep 7 20:44:12 2026 +0300

    drm/i915/dp_mst: Fix configuring FEC for a disconnected stream
    
    commit acbe9a3b60b9a6ace8ef11fe898f586251c84592 upstream.
    
    During an atomic commit after all the MST stream CRTC state is computed
    the driver ensures that the FEC is configured the same way (enabled or
    disabled) for all the streams on a given MST topology's link.
    drm_dp_mst_port_downstream_of_parent() used to determine if a stream is
    downstream of an MST port will return false if the whole topology is
    disconnected, since in that case it can't verify that the port/
    parent_port passed to it is in the given MST topology. This is a problem
    during the above FEC configuration check, since
    intel_dp_mst_check_dsc_change()->get_pipes_downstream_of_mst_ports()
    will not return all the stream CRTCs/pipes for the topology as expected.
    Since passing parent_port==NULL to get_pipes_downstream_of_mst_port()
    is meant to return all the streams for the given topology (i.e. mst_mgr)
    skip checking if an MST port is downstream of a parent port in this
    case.
    
    This fixes a problem where the FEC configuration check explained above
    failed to ensure that all streams' FEC is configured the same way if the
    topology was disconnected, leading to a FEC state mismatch error.
    
    Cc: [email protected] # v6.10+
    Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16073
    Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16384
    Reviewed-by: Luca Coelho <[email protected]>
    Signed-off-by: Imre Deak <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    (cherry picked from commit 270681fbffbba2b6ccf5b7e3c34b8b563b36167f)
    Signed-off-by: Jani Nikula <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/i915/dp_mst: Fix configuring TUs for a disconnected stream [+ + +]
Author: Imre Deak <[email protected]>
Date:   Mon Sep 7 20:44:13 2026 +0300

    drm/i915/dp_mst: Fix configuring TUs for a disconnected stream
    
    commit a443e0b8d647c1401b110d9f919d8c6cb8607260 upstream.
    
    During an atomic commit after all the MST stream CRTC state is computed
    the driver ensures that the sum of TUs of all the streams on a given MST
    topology link is within limits (63 for 8b10 and 64 for 128b132b). For a
    disconnected stream the DRM MST core's BW verification doesn't ensure
    this, because the topology state it uses for this is destroyed as soon
    as the stream (i.e. MST connector/port) is disconnected. The driver
    should keep the link state valid even for such disconnected streams, as
    userspace may disable them one-by-one only in a deferred way. Ensure the
    link's sum of TUs stays within limits in this case by simply reusing the
    maximum link BPP limit from the stream's (i.e. CRTC's) old state.
    
    The disconnection can happen either via the whole topology getting
    disconnected or via only the given stream's port getting disconnected.
    Check for both of these conditions separately, as a connector gets
    unregistered after a link disconnect event only in a deferred way.
    
    Cc: [email protected] # v6.10+
    Link: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16073
    Link: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16384
    Reviewed-by: Luca Coelho <[email protected]>
    Signed-off-by: Imre Deak <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    (cherry picked from commit ee00f8fbb2b202002ab90834e02e9ba372773a36)
    Signed-off-by: Jani Nikula <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
drm/i915/psr: Clear stale sel fetch enable bits on sel fetch disable [+ + +]
Author: Nemesa Garg <[email protected]>
Date:   Wed Sep 9 16:33:32 2026 +0530

    drm/i915/psr: Clear stale sel fetch enable bits on sel fetch disable
    
    [ Upstream commit 2777ec9852277a06ae68fee0c4f1a32783e4a999 ]
    
    Selective fetch is dropped while pipe CRC is active, and the planes keep
    their SEL_FETCH_PLANE_CTL / SEL_FETCH_CUR_CTL enable bit set in hardware
    over that. A plane disabled while selective fetch is off never gets the
    bit cleared, as the disable path is guarded by enable_psr2_sel_fetch.
    Once selective fetch comes back the hardware resumes fetching for a
    plane that is no longer enabled and keeps its DDB range reserved.
    
    Clear the bits as selective fetch is turned off instead. Atomic check
    has both the old and the new crtc state, so record the transition there
    and let the plane and cursor arm paths write the registers to 0 for that
    commit.
    
    v2: Drop the old_crtc_state->hw.active check. [Jouni]
    
    Fixes: b1f5279b5981 ("drm/i915/psr: Move plane sel fetch configuration into plane source files")
    Closes: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8739
    Assisted-by: Copilot:Claude-Opus-5
    Signed-off-by: Nemesa Garg <[email protected]>
    Reviewed-by: Jouni Högander <[email protected]>
    Signed-off-by: Suraj Kandpal <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    (cherry picked from commit a4c0e7f80429eda6990960971aebd4e4b9533cc6)
    Signed-off-by: Jani Nikula <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
drm/i915/quirks: Limit eDP rate to HBR2 on HP Pavilion Plus 14-ew1 [+ + +]
Author: Ankit Nautiyal <[email protected]>
Date:   Mon Sep 7 09:15:55 2026 +0530

    drm/i915/quirks: Limit eDP rate to HBR2 on HP Pavilion Plus 14-ew1
    
    commit 5271d81f99dd01d983d439930eb056952485e15e upstream.
    
    The eDP panel on the HP Pavilion Plus Laptop 14-ew1xxx advertises HBR3
    while leaving the TPS4 support bit clear. The output however flickers, once
    link is trained with HBR3.
    
    Until commit 8c9006283e4b ("Revert "drm/i915/dp: Reject HBR3 when sink
    doesn't support TPS4"") such sinks were capped at HBR2 by the TPS4 check
    which incidentally kept this panel stable. That check was reverted because
    other panels legitimately need HBR3 without advertising TPS4, and the
    per-machine QUIRK_EDP_LIMIT_RATE_HBR2 was introduced to handle the affected
    machines instead.
    
    Add the machine to the list of devices that need the
    QUIRK_EDP_LIMIT_RATE_HBR2.
    
    Fixes: 8c9006283e4b ("Revert "drm/i915/dp: Reject HBR3 when sink doesn't support TPS4"")
    Reported-by: Annoy Cc <[email protected]>
    Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16743
    Cc: <[email protected]> # v6.18+
    Tested-by: Annoy Cc <[email protected]>
    Signed-off-by: Ankit Nautiyal <[email protected]>
    Reviewed-by: Nemesa Garg <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    (cherry picked from commit 550b703fdbb2a2022faa75b4b11ab135241afbd9)
    Signed-off-by: Jani Nikula <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
drm/i915: fix incorrect RCU teardown order [+ + +]
Author: Christian König <[email protected]>
Date:   Thu Sep 3 13:36:21 2026 +0200

    drm/i915: fix incorrect RCU teardown order
    
    commit d2da6696e0c4e60414706e607029d0bb0330c67e upstream.
    
    i915_gem_busy_ioctl uses dma_resv_for_each_fence_unlocked() to iterate
    over the fences in an GEM object without holding a reference but only
    the RCU read side lock.
    
    What can happen here is that the GEM object is destroyed concurrently
    while i915_gem_busy_ioctl is still running. This won't free the GEM
    objects memory, but still drops all the dma_fence references.
    
    Now when dma_resv_for_each_fence_unlocked() sees a destroyed dma_fence it
    assumes that a new fence list was installed and re-starts the loop.
    
    But in the case of a destroyed GEM object a new fence list is never
    installed, only the old one freed and therefore the iteration never
    finishes resulting in an endless loop.
    
    The solution is to drop the fence references only after the RCU grace
    period.
    
    The fixes tag is not necessary the patch introducing the problem, but the
    one making it so worse that we need to address it.
    
    This problem was pointed out by Sashiko-bot.
    
    Signed-off-by: Christian König <[email protected]>
    Fixes: 912ff2ebd695 ("drm/i915: use the new iterator in i915_gem_busy_ioctl v2")
    CC: [email protected]
    Reviewed-by: Tvrtko Ursulin <[email protected]>
    Signed-off-by: Tvrtko Ursulin <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    (cherry picked from commit 5113479556025093bf8133bb2dcaa33be2d50921)
    Signed-off-by: Jani Nikula <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
drm/imagination: clamp freelist reconstruction requests [+ + +]
Author: Pengpeng Hou <[email protected]>
Date:   Sun Sep 20 11:43:29 2026 +0800

    drm/imagination: clamp freelist reconstruction requests
    
    [ Upstream commit 45585c3aa285854face65293acc95eff73063d6d ]
    
    The firmware reconstruction count controls accesses to the request's
    fixed freelist ID array and the copy into the fixed response array.
    Neither access currently bounds the count to those protocol arrays.
    
    Clamp the count to the request capacity, which is shared by the response
    layout, and use that count consistently for reconstruction and response
    publication. Keep the firmware recovery exchange instead of dropping an
    oversized request without a response, as discussed with the firmware
    maintainer.
    
    The issue was found by our static-analysis tool.
    
    Fixes: 6eedddab733b ("drm/imagination: Implement free list and HWRT create and destroy ioctls")
    Assisted-by: gpt 5
    Signed-off-by: Pengpeng Hou <[email protected]>
    Reviewed-by: Alessio Belle <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Brajesh Gupta <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

drm/imagination: Fix page count for page table for map() interface [+ + +]
Author: Brajesh Gupta <[email protected]>
Date:   Tue Sep 22 09:56:56 2026 +0530

    drm/imagination: Fix page count for page table for map() interface
    
    commit 0a8224058a5835297dcf4a46bbcd16f77a9fe424 upstream.
    
    The GPU virtual start address wasn't included in the calculation for the
    amount of page tables required for mapping a BO object in map() interface.
    It resulted in map failure later due to not enough pages at L0/L1 level.
    Update pvr_mmu_op_context_create() interface to pass device address as well
    to allow correct calculation for page table memory.
    
    If L0 tables cover 2MB (0x200000), the range defined by device address
    0x80001ff000 (general heap at 2MB - 4KB) and size 0x2000 (two 4KB pages)
    requires two L0 pages to be mapped, but without the base
    address a range of 0x2000 computes to a single L0 page which is not enough.
    
    Fixes: ff5f643de0bf ("drm/imagination: Add GEM and VM related code")
    Reviewed-by: Alexandru Dadu <[email protected]>
    Reviewed-by: Alessio Belle <[email protected]>
    Cc: [email protected]
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Brajesh Gupta <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/imagination: Propagate map failures correctly from pvr_mmu_map_sgl() [+ + +]
Author: Brajesh Gupta <[email protected]>
Date:   Tue Sep 22 09:56:55 2026 +0530

    drm/imagination: Propagate map failures correctly from pvr_mmu_map_sgl()
    
    commit 7b824c293a6b56de8285a97984c507cba56bc4c4 upstream.
    
    Map failure from pvr_mmu_map_sgl() interface was not returned correctly
    to pvr_mmu_map() interface. This resulted in pvr_mmu_map() interface to
    continue instead of returning an error to caller.
    Fix it by returning a proper error code from pvr_mmu_map_sgl() interface.
    
    Call stack for crash:
    [ 1179.286237] Unable to handle kernel NULL pointer dereference at virtual address 0000000000000008
    [ 1179.295067] Mem abort info:
    [ 1179.297877]   ESR = 0x0000000096000004
    [ 1179.301656]   EC = 0x25: DABT (current EL), IL = 32 bits
    [ 1179.306987]   SET = 0, FnV = 0
    [ 1179.310048]   EA = 0, S1PTW = 0
    [ 1179.313198]   FSC = 0x04: level 0 translation fault
    [ 1179.318096] Data abort info:
    [ 1179.320993]   ISV = 0, ISS = 0x00000004, ISS2 = 0x00000000
    [ 1179.326483]   CM = 0, WnR = 0, TnD = 0, TagAccess = 0
    [ 1179.331546]   GCS = 0, Overlay = 0, DirtyBit = 0, Xs = 0
    [ 1179.336895] user pgtable: 4k pages, 48-bit VAs, pgdp=000000009822a000
    [ 1179.343402] [0000000000000008] pgd=0000000000000000, p4d=0000000000000000
    [ 1179.350243] Internal error: Oops: 0000000096000004 [#2]  SMP
    [ 1179.355908] Modules linked in: powervr gpu_sched drm_shmem_helper drm_gpuvm drm_exec xhci_plat_hcd xhci_hcd dwc3 usbcore usb_common snd_soc_simple_card snd_soc_simple_card_utils dwc3_am62 at24 sa2ul sha512 libsha512 sha256 authenc sch_fq_codel fuse dm_mod ipv6
    [ 1179.378992] CPU: 1 UID: 1000 PID: 680 Comm: deqp-vk Tainted: G      D             6.17.0 #1 PREEMPT
    [ 1179.388120] Tainted: [D]=DIE
    [ 1179.390994] Hardware name: Texas Instruments AM625 SK (DT)
    [ 1179.396467] pstate: 00000005 (nzcv daif -PAN -UAO -TCO -DIT -SSBS BTYPE=--)
    [ 1179.403415] pc : pvr_mmu_op_context_unmap_curr_page+0x6c/0x134 [powervr]
    [ 1179.410140] lr : pvr_mmu_op_context_unmap_curr_page+0x58/0x134 [powervr]
    [ 1179.416848] sp : ffff8000839ab8c0
    [ 1179.420153] x29: ffff8000839ab8c0 x28: 0000000000000001 x27: 000000008f386000
    [ 1179.427283] x26: ffff000016d1df98 x25: 0000000000247000 x24: 00000000000001e6
    [ 1179.434413] x23: 0000000000000002 x22: 000000000000ffff x21: 0000000000000247
    [ 1179.441540] x20: 0000000000000245 x19: ffff000016d1df60 x18: 0000000000000002
    [ 1179.448668] x17: 0000000000000000 x16: 0000000000000000 x15: 0000000000000001
    [ 1179.455793] x14: 0000000000060810 x13: ffff80007fffffff x12: ffff000004190480
    [ 1179.462921] x11: ffff8000853f7000 x10: ffff8000811ae000 x9 : ffff0000041900b8
    [ 1179.470051] x8 : 0000000000000000 x7 : 00000000990c4001 x6 : 0000000000000007
    [ 1179.477177] x5 : ffff000016d1df60 x4 : 0000000000000000 x3 : ffff00000a7d8000
    [ 1179.484306] x2 : 00000000000001ff x1 : 0000000000000000 x0 : 0000000000000000
    [ 1179.491433] Call trace:
    [ 1179.493872]  pvr_mmu_op_context_unmap_curr_page+0x6c/0x134 [powervr] (P)
    [ 1179.500582]  pvr_mmu_map+0x31c/0x388 [powervr]
    [ 1179.505027]  pvr_vm_gpuva_map+0x40/0x88 [powervr]
    [ 1179.509732]  __drm_gpuvm_sm_map+0x250/0x44c [drm_gpuvm]
    [ 1179.514952]  drm_gpuvm_sm_map+0x48/0x5c [drm_gpuvm]
    [ 1179.519822]  pvr_vm_bind_op_exec+0x64/0x70 [powervr]
    [ 1179.524785]  pvr_vm_map+0x1f8/0x2a8 [powervr]
    [ 1179.529142]  pvr_ioctl_vm_map+0x12c/0x188 [powervr]
    [ 1179.534018]  drm_ioctl_kernel+0xb8/0x128
    [ 1179.537941]  drm_ioctl+0x21c/0x4ec
    [ 1179.541337]  __arm64_sys_ioctl+0xac/0x108
    [ 1179.545344]  invoke_syscall+0x44/0x100
    [ 1179.549091]  el0_svc_common.constprop.0+0x40/0xe0
    [ 1179.553790]  do_el0_svc+0x1c/0x28
    [ 1179.557106]  el0_svc+0x34/0xf0
    [ 1179.560159]  el0t_64_sync_handler+0xd0/0xe4
    [ 1179.564334]  el0t_64_sync+0x198/0x19c
    [ 1179.567996] Code: 54000300 35000360 f9402261 79409a62 (f9400421)
    [ 1179.574081] ---[ end trace 0000000000000000 ]---
    
    Fixes: ff5f643de0bf ("drm/imagination: Add GEM and VM related code")
    Reviewed-by: Alexandru Dadu <[email protected]>
    Reviewed-by: Alessio Belle <[email protected]>
    Cc: [email protected]
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Brajesh Gupta <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
drm/nouveau/clk: don't clobber reclock status when restoring volt/fan [+ + +]
Author: Francesco Magazzu <[email protected]>
Date:   Fri Sep 18 15:16:20 2026 +0200

    drm/nouveau/clk: don't clobber reclock status when restoring volt/fan
    
    [ Upstream commit e5cccdafc855cd5f96f4b51d38114a0360b075d7 ]
    
    nvkm_cstate_prog() reuses 'ret' for the voltage and fan-speed restore
    calls it makes after reprogramming the clocks.  Those calls almost always
    succeed, so the status of the reclock itself is overwritten and the
    function reports success even when clk->func->calc() or clk->func->prog()
    failed.  The converse is also true: a successful reclock is reported as an
    error if the final restore call fails, even though that failure is only
    logged and otherwise ignored.
    
    The only consumer of the return value is the error message in
    nvkm_pstate_work(), so in practice a failing reclock is simply never
    reported.  Nothing else changes, but a function that returns success on
    failure is a trap for the next caller.
    
    Keep the calc/prog status in 'ret' and use a separate local for the
    restore calls.
    
    Fixes: 3eca809b3c05 ("drm/nouveau/clk: cosmetic changes")
    Signed-off-by: Francesco Magazzu <[email protected]>
    Reviewed-by: Lyude Paul <[email protected]>
    Signed-off-by: Lyude Paul <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

drm/nouveau/clk: fix list cursor use after loop in nvkm_clk_ustate_update [+ + +]
Author: Dan Carpenter <[email protected]>
Date:   Fri Sep 18 15:16:17 2026 +0200

    drm/nouveau/clk: fix list cursor use after loop in nvkm_clk_ustate_update
    
    [ Upstream commit aff09d9e37e02dc60bde79035ac15b136d602259 ]
    
    If list_for_each_entry() exits without hitting a break then "pstate" is
    not a valid pstate pointer.  Introduce a "found" variable instead.
    
    The check is reachable from userspace: nvkm_clk_ustate_update() takes the
    pstate id straight from the 'pstate' debugfs file, so requesting an id
    that is not in clk->states - or any id at all when the perf tables are
    broken and the list is empty - makes the pstate->pstate != req test
    dereference the list head cast to a struct nvkm_pstate, which is an
    out-of-bounds read.
    
    Fixes: 7c8565220697 ("drm/nouveau/clk: implement power state and engine clock control in core")
    Signed-off-by: Dan Carpenter <[email protected]>
    [Francesco: rebased on drm-misc-next, expanded the commit message]
    Signed-off-by: Francesco Magazzu <[email protected]>
    Reviewed-by: Lyude Paul <[email protected]>
    Signed-off-by: Lyude Paul <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

 
drm/nouveau/disp: don't reject HDMI config on cards without SCDC [+ + +]
Author: Giuseppe Ranieri <[email protected]>
Date:   Thu Sep 17 21:51:14 2026 +0000

    drm/nouveau/disp: don't reject HDMI config on cards without SCDC
    
    [ Upstream commit 1717fcc5be575d4768279148ae9465a8b13d4339 ]
    
    nv50_hdmi_enable() passes the sink's SCDC capability from its EDID
    straight through to nvif_outp_hdmi(). On pre-Maxwell-2 cards there is no
    hdmi->scdc callback, so nvkm_uoutp_mthd_hdmi() rejects the whole
    configuration with -EINVAL, and nv50_hdmi_enable() returns before
    hdmi->ctrl() runs and before the AVI and VSI infoframes are sent.
    
    The result on such a card driving an SCDC-capable HDMI 2.0 sink is that
    HDMI audio silently stops working. Video is unaffected, and nothing is
    logged, which makes the failure hard to attribute.
    
    SCDC is optional, and the hdmi->scdc() call further down is already
    guarded against a missing callback. Requesting it on a card that cannot
    do it need not invalidate the rest of the HDMI configuration, so drop
    that term from the condition and let the existing guard skip SCDC alone.
    
    Fixes: 6c6abab20b99 ("drm/nouveau/disp: add output hdmi config method")
    Signed-off-by: Giuseppe Ranieri <[email protected]>
    Co-Authored-By: Tano Dzhinski <[email protected]>
    Signed-off-by: Tano Dzhinski <[email protected]>
    Tested-by: Tano Dzhinski <[email protected]>
    Reviewed-by: Lyude Paul <[email protected]>
    Signed-off-by: Lyude Paul <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

 
drm/nouveau/dmem: pin VRAM for the whole registered range [+ + +]
Author: Junrui Luo <[email protected]>
Date:   Mon Aug 17 14:50:40 2026 +0800

    drm/nouveau/dmem: pin VRAM for the whole registered range
    
    commit 64ca4cdd1031206424e6455f0ae4fb0560d8f46a upstream.
    
    Commit c32287471077 ("gpu/drm/nouveau: enable THP support for GPU memory
    migration") grew the device-private region that
    nouveau_dmem_chunk_alloc() registers from DMEM_CHUNK_SIZE to
    DMEM_CHUNK_SIZE * NR_CHUNKS, but left the VRAM buffer object backing that
    region at DMEM_CHUNK_SIZE.
    
    nouveau_dmem_page_addr() returns chunk->bo->offset plus the page's offset
    within the registered region, so every page past the first chunk resolves
    to VRAM outside the buffer object.
    
    Size the buffer object to the region it backs.
    
    Fixes: c32287471077 ("gpu/drm/nouveau: enable THP support for GPU memory migration")
    Reported-by: Yuhao Jiang <[email protected]>
    Assisted-by: LLM
    Cc: [email protected]
    Signed-off-by: Junrui Luo <[email protected]>
    Reviewed-by: Lyude Paul <[email protected]>
    Signed-off-by: Lyude Paul <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
drm/nouveau/uvmm: fix UAF in nouveau_uvmm_sm when BO is in TTM_PL_SYSTEM [+ + +]
Author: Peiyang He <[email protected]>
Date:   Mon Sep 7 13:22:04 2026 +0800

    drm/nouveau/uvmm: fix UAF in nouveau_uvmm_sm when BO is in TTM_PL_SYSTEM
    
    commit 3359a372efb6d585c97019ee1b7f1874442bcebe upstream.
    
    nouveau_uvmm_sm() calls op_map(), which passes bo->resource through
    nouveau_mem() to nouveau_uvma_map(). nouveau_uvmm_vmm_map() then reads
    mem->mem.type.
    
    But this is only valid when bo->resource is backed by struct nouveau_mem,
    as is the case for VRAM and TT resources. If the BO is left in
    TTM_PL_SYSTEM, bo->resource is only a struct ttm_resource. Treating it
    as struct nouveau_mem makes the mem->mem.type read past the end of the
    resource, causing a KASAN: slab-use-after-free Read in nouveau_uvmm_sm
    report:
    
    BUG: KASAN: slab-use-after-free in nouveau_uvmm_vmm_map drivers/gpu/drm/nouveau/nouveau_uvmm.c:152 [inline]
    BUG: KASAN: slab-use-after-free in nouveau_uvma_map drivers/gpu/drm/nouveau/nouveau_uvmm.c:199 [inline]
    BUG: KASAN: slab-use-after-free in op_map drivers/gpu/drm/nouveau/nouveau_uvmm.c:849 [inline]
    BUG: KASAN: slab-use-after-free in nouveau_uvmm_sm.constprop.0+0x6ab/0x900 drivers/gpu/drm/nouveau/nouveau_uvmm.c:903
    Read of size 1 at addr ffff888127d3e3a0 by task kworker/0:1/11
    
    CPU: 0 UID: 0 PID: 11 Comm: kworker/0:1 Not tainted 7.2.0 #5 PREEMPT(lazy)
    Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
    Workqueue: nouveau_sched_wq_2224 drm_sched_run_job_work
    Call Trace:
     <TASK>
     __dump_stack lib/dump_stack.c:94 [inline]
     dump_stack_lvl+0x95/0xe0 lib/dump_stack.c:120
     print_address_description mm/kasan/report.c:378 [inline]
     print_report+0xcb/0x5a0 mm/kasan/report.c:482
     kasan_report+0xca/0x100 mm/kasan/report.c:595
     nouveau_uvmm_vmm_map drivers/gpu/drm/nouveau/nouveau_uvmm.c:152 [inline]
     nouveau_uvma_map drivers/gpu/drm/nouveau/nouveau_uvmm.c:199 [inline]
     op_map drivers/gpu/drm/nouveau/nouveau_uvmm.c:849 [inline]
     nouveau_uvmm_sm.constprop.0+0x6ab/0x900 drivers/gpu/drm/nouveau/nouveau_uvmm.c:903
     nouveau_uvmm_sm_unmap drivers/gpu/drm/nouveau/nouveau_uvmm.c:932 [inline]
     nouveau_uvmm_bind_job_run+0xd6/0x250 drivers/gpu/drm/nouveau/nouveau_uvmm.c:1532
     nouveau_job_run drivers/gpu/drm/nouveau/nouveau_sched.c:350 [inline]
     nouveau_sched_run_job+0x62/0xd0 drivers/gpu/drm/nouveau/nouveau_sched.c:364
     drm_sched_run_job_work+0x356/0xa10 drivers/gpu/drm/scheduler/sched_main.c:1061
     process_one_work+0x8a5/0x1900 kernel/workqueue.c:3322
     process_scheduled_works kernel/workqueue.c:3405 [inline]
     worker_thread+0x5dd/0xd80 kernel/workqueue.c:3486
     kthread+0x31d/0x420 kernel/kthread.c:436
     ret_from_fork+0x662/0x940 arch/x86/kernel/process.c:158
     ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
     </TASK>
    
    Allocated by task 2224 on cpu 0 at 66.550027s:
     kasan_save_stack+0x24/0x50 mm/kasan/common.c:57
     kasan_save_track+0x17/0x60 mm/kasan/common.c:78
     poison_kmalloc_redzone mm/kasan/common.c:398 [inline]
     __kasan_kmalloc+0xaa/0xb0 mm/kasan/common.c:415
     kasan_kmalloc include/linux/kasan.h:263 [inline]
     __do_kmalloc_node mm/slub.c:5334 [inline]
     __kmalloc_noprof+0x304/0x7c0 mm/slub.c:5359
     _kmalloc_noprof include/linux/slab.h:992 [inline]
     dma_resv_list_alloc+0x27/0x90 drivers/dma-buf/dma-resv.c:106
     dma_resv_reserve_fences+0x60e/0xa30 drivers/dma-buf/dma-resv.c:205
     ttm_bo_alloc_resource+0x12c/0xbd0 drivers/gpu/drm/ttm/ttm_bo.c:721
     ttm_bo_validate+0x1bc/0x4a0 drivers/gpu/drm/ttm/ttm_bo.c:856
     ttm_bo_init_reserved+0x2c3/0x570 drivers/gpu/drm/ttm/ttm_bo.c:970
     nouveau_bo_init+0x159/0x2c0 drivers/gpu/drm/nouveau/nouveau_bo.c:359
     nouveau_gem_new+0x234/0x5f0 drivers/gpu/drm/nouveau/nouveau_gem.c:272
     nouveau_gem_ioctl_new+0x1eb/0x420 drivers/gpu/drm/nouveau/nouveau_gem.c:352
     drm_ioctl_kernel+0x192/0x350 drivers/gpu/drm/drm_ioctl.c:817
     drm_ioctl+0x4f8/0xb40 drivers/gpu/drm/drm_ioctl.c:914
     nouveau_drm_ioctl+0xea/0x2c0 drivers/gpu/drm/nouveau/nouveau_drm.c:1338
     vfs_ioctl fs/ioctl.c:51 [inline]
     __do_sys_ioctl fs/ioctl.c:597 [inline]
     __se_sys_ioctl fs/ioctl.c:583 [inline]
     __x64_sys_ioctl+0x180/0x1d0 fs/ioctl.c:583
     do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
     do_syscall_64+0x115/0x690 arch/x86/entry/syscall_64.c:94
     entry_SYSCALL_64_after_hwframe+0x77/0x7f
    
    Freed by task 2223 on cpu 0 at 66.554063s:
     kasan_save_stack+0x24/0x50 mm/kasan/common.c:57
     kasan_save_track+0x17/0x60 mm/kasan/common.c:78
     kasan_save_free_info+0x3b/0x60 mm/kasan/generic.c:584
     poison_slab_object mm/kasan/common.c:253 [inline]
     __kasan_slab_free+0x61/0x80 mm/kasan/common.c:285
     kasan_slab_free include/linux/kasan.h:235 [inline]
     slab_free_hook mm/slub.c:2677 [inline]
     __rcu_free_sheaf_prepare+0xb6/0x2e0 mm/slub.c:2928
     rcu_free_sheaf+0x1b/0x120 mm/slub.c:5978
     rcu_do_batch kernel/rcu/tree.c:2645 [inline]
     rcu_core+0x521/0x1490 kernel/rcu/tree.c:2897
     handle_softirqs+0x1b1/0x8a0 kernel/softirq.c:622
     __do_softirq kernel/softirq.c:656 [inline]
     invoke_softirq kernel/softirq.c:496 [inline]
     __irq_exit_rcu+0x137/0x1c0 kernel/softirq.c:735
     irq_exit_rcu+0x9/0x20 kernel/softirq.c:752
     instr_sysvec_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1062 [inline]
     sysvec_apic_timer_interrupt+0x70/0x80 arch/x86/kernel/apic/apic.c:1062
     asm_sysvec_apic_timer_interrupt+0x1a/0x20 arch/x86/include/asm/idtentry.h:674
    
    The buggy address belongs to the object at ffff888127d3e380
     which belongs to the cache kmalloc-96 of size 96
    The buggy address is located 32 bytes inside of
     freed 96-byte region [ffff888127d3e380, ffff888127d3e3e0)
    
    The buggy address belongs to the physical page:
    page: refcount:0 mapcount:0 mapping:0000000000000000 index:0x0 pfn:0x127d3e
    flags: 0x200000000000000(node=0|zone=2)
    page_type: f5(slab)
    raw: 0200000000000000 ffff888100041280 dead000000000122 0000000000000000
    raw: 0000000000000000 0000000000200020 00000000f5000000 0000000000000000
    page dumped because: kasan: bad access detected
    
    Memory state around the buggy address:
     ffff888127d3e280: 00 00 00 00 00 00 00 00 00 fc fc fc fc fc fc fc
     ffff888127d3e300: 00 00 00 00 00 00 00 00 00 00 00 fc fc fc fc fc
    >ffff888127d3e380: fa fb fb fb fb fb fb fb fb fb fb fb fc fc fc fc
                                   ^
     ffff888127d3e400: fa fb fb fb fb fb fb fb fb fb fb fb fc fc fc fc
     ffff888127d3e480: 00 00 00 00 00 00 00 00 00 00 fc fc fc fc fc fc
    
    Fix by resetting the placement to the BO's valid domains before
    calling nouveau_bo_validate(), matching the handling in
    nouveau_uvmm_bo_validate(), so map jobs do not run for SYSTEM resources;
    Reject BO that cannot reside in VRAM or GART;
    Also skip op_map() when the GPUVA has been invalidated, matching the
    handling in the unmap and remap paths.
    
    Found when fuzzing the nouveau driver with a modified Syzkaller.
    
    Fixes: b88baab82871 ("drm/nouveau: implement new VM_BIND uAPI")
    Cc: [email protected]
    Signed-off-by: Peiyang He <[email protected]>
    Assisted-by: Codex:gpt-5.5
    Reviewed-by: Lyude Paul <[email protected]>
    Signed-off-by: Lyude Paul <[email protected]>
    Link: https://patch.msgid.link/0D77BEC410CE0129+20260907052204.1431488-1-peiyang_he@smail.nju.edu.cn
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
drm/nouveau: don't bump pin count on failed re-pin in nouveau_bo_pin_locked() [+ + +]
Author: Peiyang He <[email protected]>
Date:   Fri Sep 18 10:53:12 2026 +0800

    drm/nouveau: don't bump pin count on failed re-pin in nouveau_bo_pin_locked()
    
    commit 6a6870d3077faa501ca97760057ddca22b68418d upstream.
    
    nouveau_bo_pin_locked() checks whether an already pinned BO is in a
    memory domain compatible with a new pin request. When the domains are
    incompatible, it sets -EBUSY but still calls ttm_bo_pin() before
    returning.
    
    Callers treat a failed nouveau_bo_pin() as not having acquired a new pin,
    so the extra pin count is never decreased by a matching unpin.
    This triggers the warning in ttm_bo_release():
    
            WARN_ON_ONCE(bo->pin_count);
    
    Found when fuzzing the nouveau driver with a modified Syzkaller:
    
            WARNING: drivers/gpu/drm/ttm/ttm_bo.c:256 at ttm_bo_release+0x827/0x9e0 drivers/gpu/drm/ttm/ttm_bo.c:256, CPU#1: syz.3.24/2212
            Modules linked in:
            CPU: 1 UID: 0 PID: 2212 Comm: syz.3.24 Not tainted 7.2.0 #24 PREEMPT(lazy)
            nouveau 0000:01:00.0: gsp:msg fn:103 len:0x40/0x20 res:0x19 resp:0x19
            Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
            RIP: 0010:ttm_bo_release+0x827/0x9e0 drivers/gpu/drm/ttm/ttm_bo.c:256
            Code: 02 00 0f 85 51 01 00 00 48 8b 7b 08 e8 d2 20 01 00 e9 80 fd ff ff e8 d8 15 c0 fe 90 0f 0b 90 e9 e1 f8 ff ff e8 ca 15 c0 fe 90 <0f> 0b 90 e9 a4 f8 ff ff e8 bc 15 c0 fe be 03 00 00 00 4c 89 e7 e8
            msg: 00000000: 05 00 d0 c1 04 00 f0 f1 01 30 00 00 2d 90 00 00  .........0..-...
            RSP: 0018:ffffc9000f5cf710 EFLAGS: 00010293
            RAX: 0000000000000000 RBX: ffff888018e5d2a8 RCX: ffffffff82bb1b36
            RDX: ffff888017b68000 RSI: 0000000000000004 RDI: ffff888018e5d2a8
            msg: 00000010: 19 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00  ................
            RBP: ffff88801261c720 R08: 0000000000000001 R09: ffffed10031cba55
            R10: ffff888018e5d2ab R11: 00000000000000f3 R12: ffff888018e5d290
            R13: ffff888018e5d2d4 R14: ffff88801b219c18 R15: dffffc0000000000
            FS:  0000000000000000(0000) GS:ffff8880e0f6f000(0000) knlGS:0000000000000000
            CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
            CR2: 0000001b31223ffc CR3: 0000000028e00005 CR4: 0000000000770ef0
            PKRU: 80000000
            Call Trace:
            <TASK>
            kref_put include/linux/kref.h:65 [inline]
            ttm_bo_put drivers/gpu/drm/ttm/ttm_bo.c:325 [inline]
            ttm_bo_fini+0x55/0x80 drivers/gpu/drm/ttm/ttm_bo.c:330
            nouveau_gem_object_del+0xb2/0x1b0 drivers/gpu/drm/nouveau/nouveau_gem.c:90
            drm_gem_object_free+0x5f/0x90 drivers/gpu/drm/drm_gem.c:1165
            kref_put include/linux/kref.h:65 [inline]
            __drm_gem_object_put include/drm/drm_gem.h:562 [inline]
            drm_gem_object_put include/drm/drm_gem.h:575 [inline]
            nouveau_abi16_chan_fini.constprop.0+0x44f/0x5a0 drivers/gpu/drm/nouveau/nouveau_abi16.c:195
            nouveau 0000:01:00.0: syz.2.23[2209]: Unknown handle 0x00000000
            nouveau_abi16_fini+0x1d0/0x340 drivers/gpu/drm/nouveau/nouveau_abi16.c:225
            nouveau_drm_postclose+0x18b/0x3e0 drivers/gpu/drm/nouveau/nouveau_drm.c:1284
            nouveau 0000:01:00.0: syz.2.23[2209]: validate_init
            drm_file_free.part.0+0x6d6/0xb60 drivers/gpu/drm/drm_file.c:267
            drm_file_free drivers/gpu/drm/drm_file.c:237 [inline]
            drm_close_helper.isra.0+0x11a/0x160 drivers/gpu/drm/drm_file.c:290
            drm_release+0x1ab/0x330 drivers/gpu/drm/drm_file.c:438
            __fput+0x39c/0xa60 fs/file_table.c:512
            nouveau 0000:01:00.0: syz.2.23[2209]: validate: -2
            task_work_run+0x15a/0x230 kernel/task_work.c:233
            exit_task_work include/linux/task_work.h:40 [inline]
            do_exit+0x82b/0x25a0 kernel/exit.c:1009
            do_group_exit+0xc2/0x280 kernel/exit.c:1152
            get_signal+0x1d6e/0x1f30 kernel/signal.c:3046
            arch_do_signal_or_restart+0x7d/0x6e0 arch/x86/kernel/signal.c:337
            __exit_to_user_mode_loop kernel/entry/common.c:66 [inline]
            exit_to_user_mode_loop+0xdf/0x440 kernel/entry/common.c:101
            __exit_to_user_mode_prepare include/linux/irq-entry-common.h:207 [inline]
            syscall_exit_to_user_mode_prepare include/linux/irq-entry-common.h:230 [inline]
            syscall_exit_to_user_mode include/linux/entry-common.h:318 [inline]
            do_syscall_64+0x4f8/0x690 arch/x86/entry/syscall_64.c:100
            entry_SYSCALL_64_after_hwframe+0x77/0x7f
            RIP: 0033:0x7f12bac8594d
            Code: Unable to access opcode bytes at 0x7f12bac85923.
            RSP: 002b:00007f12b96e70d8 EFLAGS: 00000246 ORIG_RAX: 00000000000000ca
            RAX: 0000000000000001 RBX: 00007f12baf15fa8 RCX: 00007f12bac8594d
            RDX: 00000000000f4240 RSI: 0000000000000081 RDI: 00007f12baf15fac
            RBP: 00007f12baf15fa0 R08: 00007f12baee8000 R09: 0000000000000000
            R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
            R13: 00007f12baf16038 R14: 0000000000000006 R15: 00007ffe2ed394b0
            </TASK>
            irq event stamp: 47867
            hardirqs last  enabled at (47883): [<ffffffff815cafc6>] __up_console_sem+0x66/0x70 kernel/printk/printk.c:347
            hardirqs last disabled at (47892): [<ffffffff815cafab>] __up_console_sem+0x4b/0x70 kernel/printk/printk.c:345
            softirqs last  enabled at (47880): [<ffffffff81434277>] __do_softirq kernel/softirq.c:656 [inline]
            softirqs last  enabled at (47880): [<ffffffff81434277>] invoke_softirq kernel/softirq.c:496 [inline]
            softirqs last  enabled at (47880): [<ffffffff81434277>] __irq_exit_rcu+0x137/0x1c0 kernel/softirq.c:735
            softirqs last disabled at (47875): [<ffffffff81434277>] __do_softirq kernel/softirq.c:656 [inline]
            softirqs last disabled at (47875): [<ffffffff81434277>] invoke_softirq kernel/softirq.c:496 [inline]
            softirqs last disabled at (47875): [<ffffffff81434277>] __irq_exit_rcu+0x137/0x1c0 kernel/softirq.c:735
    
    Fix by calling ttm_bo_pin() only when the existing placement is compatible
    with the new pin request. This matches the correct behavior in other DRM
    drivers such as amdgpu_bo_pin() in amdgpu.
    
    Cc: [email protected]
    Fixes: ad76b3f7c7a0 ("drm/nouveau: teach nouveau_bo_pin() how to force a contig vram allocation")
    Signed-off-by: Peiyang He <[email protected]>
    Assisted-by: LLM
    Reviewed-by: Lyude Paul <[email protected]>
    Signed-off-by: Lyude Paul <[email protected]>
    Link: https://patch.msgid.link/EACEF2F4E098413F+20260918025312.2814889-1-peiyang_he@smail.nju.edu.cn
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/nouveau: fix autosuspend cleanup during teardown [+ + +]
Author: Guangshuo Li <[email protected]>
Date:   Sat Aug 8 21:41:37 2026 +0800

    drm/nouveau: fix autosuspend cleanup during teardown
    
    commit fefd9480ec361969f1a836df46326a1801062c26 upstream.
    
    nouveau_drm_device_init() calls pm_runtime_use_autosuspend(), but
    nouveau_drm_device_fini() does not call the matching
    pm_runtime_dont_use_autosuspend().
    
    If the autosuspend delay is set to a negative value while autosuspend
    is enabled, the runtime PM core increments usage_count to prevent
    runtime suspend. Without calling pm_runtime_dont_use_autosuspend()
    during teardown, this reference is not dropped and usage_count remains
    unbalanced.
    
    The documentation for pm_runtime_use_autosuspend() also notes that it
    is important to undo it with pm_runtime_dont_use_autosuspend() at
    driver exit time, unless runtime PM was initially enabled with
    devm_pm_runtime_enable().
    
    Add the missing pm_runtime_dont_use_autosuspend() call to the common
    device teardown path.
    
    This issue was found by manual code inspection.
    
    Fixes: 5addcf0a5f0f ("nouveau: add runtime PM support (v0.9)")
    Cc: [email protected]
    Signed-off-by: Guangshuo Li <[email protected]>
    Reviewed-by: Lyude Paul <[email protected]>
    Signed-off-by: Lyude Paul <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/nouveau: Fix bridge reference leak in nv1a_ram_new() [+ + +]
Author: Wentao Liang <[email protected]>
Date:   Wed Sep 16 18:00:36 2026 +0000

    drm/nouveau: Fix bridge reference leak in nv1a_ram_new()
    
    commit 67b4411538c8341692548429d43256f25be99f7a upstream.
    
    pci_get_domain_bus_and_slot() takes a reference to the PCI device,
    which is never released once the memory size has been read from its
    config space.  Drop the reference before returning.
    
    Fixes: 2fa6d6cdaf283c05 ("drm/nouveau: deprecate pci_get_bus_and_slot()")
    Cc: [email protected]
    Signed-off-by: Wentao Liang <[email protected]>
    Reviewed-by: Lyude Paul <[email protected]>
    Signed-off-by: Lyude Paul <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/nouveau: fix double-free in nvif_vmm_dtor [+ + +]
Author: Peiyang He <[email protected]>
Date:   Wed Sep 16 18:31:38 2026 +0800

    drm/nouveau: fix double-free in nvif_vmm_dtor
    
    commit 97077ac87afe9e91ec074ef0be64454e7ccbf344 upstream.
    
    On failure, nouveau_cli_init() calls nouveau_cli_fini() to tear
    the client down. Then, nouveau_drm_open() also enters into its
    cleanup path and calls nouveau_cli_fini() AGAIN. nouveau_cli_fini()
    calls nouveau_vmm_fini():
    
            void
            nouveau_vmm_fini(struct nouveau_vmm *vmm)
            {
                    nouveau_svmm_fini(&vmm->svmm);
                    nvif_vmm_dtor(&vmm->vmm);
                    vmm->cli = NULL;
            }
    
    Inside nvif_vmm_dtor(), vmm->page is freed unconditionally:
    
            void
            nvif_vmm_dtor(struct nvif_vmm *vmm)
            {
                    kfree(vmm->page);
                    nvif_object_dtor(&vmm->object);
            }
    
    vmm->page is never cleared after being freed, so the second call of
    nvif_vmm_dtor() will cause a double-free.
    
    Found by fuzzing the nouveau driver with a modified Syzkaller:
    
            BUG: KASAN: double-free in nvif_vmm_dtor+0x31/0x50 drivers/gpu/drm/nouveau/nvif/vmm.c:194
            Free of addr ffff888010fcdc30 by task syz.0.173/2567
    
            CPU: 1 UID: 0 PID: 2567 Comm: syz.0.173 Not tainted 7.2.0 #24 PREEMPT(lazy)
            Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
            Call Trace:
            <TASK>
            __dump_stack lib/dump_stack.c:94 [inline]
            dump_stack_lvl+0x95/0xe0 lib/dump_stack.c:120
            print_address_description mm/kasan/report.c:378 [inline]
            print_report+0xcb/0x5a0 mm/kasan/report.c:482
            kasan_report_invalid_free+0xaa/0xd0 mm/kasan/report.c:557
            check_slab_allocation+0xe4/0x110 mm/kasan/common.c:235
            kasan_slab_pre_free include/linux/kasan.h:199 [inline]
            slab_free_hook mm/slub.c:2622 [inline]
            slab_free mm/slub.c:6377 [inline]
            kfree+0x192/0x590 mm/slub.c:6692
            nvif_vmm_dtor+0x31/0x50 drivers/gpu/drm/nouveau/nvif/vmm.c:194
            nouveau_vmm_fini+0x16/0x50 drivers/gpu/drm/nouveau/nouveau_vmm.c:127
            nouveau_cli_fini+0x10e/0x210 drivers/gpu/drm/nouveau/nouveau_drm.c:225
            nouveau_drm_open+0x24e/0x740 drivers/gpu/drm/nouveau/nouveau_drm.c:1255
            drm_file_alloc+0x5f2/0xad0 drivers/gpu/drm/drm_file.c:176
            drm_open_helper+0x1d7/0x4a0 drivers/gpu/drm/drm_file.c:335
            drm_open+0x190/0x3d0 drivers/gpu/drm/drm_file.c:388
            drm_stub_open+0x1f2/0x460 drivers/gpu/drm/drm_drv.c:1211
            chrdev_open+0x21c/0x660 fs/char_dev.c:411
            do_dentry_open+0x59d/0x12b0 fs/open.c:947
            vfs_open+0x82/0x390 fs/open.c:1052
            do_open fs/namei.c:4700 [inline]
            path_openat+0x2345/0x3420 fs/namei.c:4863
            do_file_open+0x207/0x460 fs/namei.c:4892
            do_sys_openat2+0xd1/0x1d0 fs/open.c:1368
            do_sys_open fs/open.c:1374 [inline]
            __do_sys_openat fs/open.c:1390 [inline]
            __se_sys_openat fs/open.c:1385 [inline]
            __x64_sys_openat+0x144/0x200 fs/open.c:1385
            do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
            do_syscall_64+0x115/0x690 arch/x86/entry/syscall_64.c:94
            entry_SYSCALL_64_after_hwframe+0x77/0x7f
            RIP: 0033:0x7fc6d687594d
            Code: ff c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 b0 ff ff ff f7 d8 64 89 01 48
            RSP: 002b:00007fc6d5295008 EFLAGS: 00000246 ORIG_RAX: 0000000000000101
            RAX: ffffffffffffffda RBX: 00007fc6d6b06180 RCX: 00007fc6d687594d
            RDX: 0000000000022501 RSI: 0000200000000000 RDI: ffffffffffffff9c
            RBP: 00007fc6d691c303 R08: 0000000000000000 R09: 0000000000000000
            R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
            R13: 00007fc6d6b06218 R14: 00007fc6d6b06180 R15: 00007ffd9451d760
            </TASK>
    
            Allocated by task 2567 on cpu 1 at 163.593900s:
            kasan_save_stack+0x24/0x50 mm/kasan/common.c:57
            kasan_save_track+0x17/0x60 mm/kasan/common.c:78
            poison_kmalloc_redzone mm/kasan/common.c:398 [inline]
            __kasan_kmalloc+0xaa/0xb0 mm/kasan/common.c:415
            kasan_kmalloc include/linux/kasan.h:263 [inline]
            __do_kmalloc_node mm/slub.c:5334 [inline]
            __kmalloc_noprof+0x304/0x7c0 mm/slub.c:5359
            _kmalloc_noprof include/linux/slab.h:992 [inline]
            nvif_vmm_ctor+0x3c0/0x7e0 drivers/gpu/drm/nouveau/nvif/vmm.c:237
            nouveau_vmm_init+0x40/0x90 drivers/gpu/drm/nouveau/nouveau_vmm.c:134
            nouveau_cli_init+0x7b9/0xe10 drivers/gpu/drm/nouveau/nouveau_drm.c:293
            nouveau_drm_open+0x236/0x740 drivers/gpu/drm/nouveau/nouveau_drm.c:1243
            drm_file_alloc+0x5f2/0xad0 drivers/gpu/drm/drm_file.c:176
            drm_open_helper+0x1d7/0x4a0 drivers/gpu/drm/drm_file.c:335
            drm_open+0x190/0x3d0 drivers/gpu/drm/drm_file.c:388
            drm_stub_open+0x1f2/0x460 drivers/gpu/drm/drm_drv.c:1211
            chrdev_open+0x21c/0x660 fs/char_dev.c:411
            do_dentry_open+0x59d/0x12b0 fs/open.c:947
            vfs_open+0x82/0x390 fs/open.c:1052
            do_open fs/namei.c:4700 [inline]
            path_openat+0x2345/0x3420 fs/namei.c:4863
            do_file_open+0x207/0x460 fs/namei.c:4892
            do_sys_openat2+0xd1/0x1d0 fs/open.c:1368
            do_sys_open fs/open.c:1374 [inline]
            __do_sys_openat fs/open.c:1390 [inline]
            __se_sys_openat fs/open.c:1385 [inline]
            __x64_sys_openat+0x144/0x200 fs/open.c:1385
            do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
            do_syscall_64+0x115/0x690 arch/x86/entry/syscall_64.c:94
            entry_SYSCALL_64_after_hwframe+0x77/0x7f
    
            Freed by task 2567 on cpu 1 at 163.601355s:
            kasan_save_stack+0x24/0x50 mm/kasan/common.c:57
            kasan_save_track+0x17/0x60 mm/kasan/common.c:78
            kasan_save_free_info+0x3b/0x60 mm/kasan/generic.c:584
            poison_slab_object mm/kasan/common.c:253 [inline]
            __kasan_slab_free+0x61/0x80 mm/kasan/common.c:285
            kasan_slab_free include/linux/kasan.h:235 [inline]
            slab_free_hook mm/slub.c:2677 [inline]
            slab_free mm/slub.c:6377 [inline]
            kfree+0x383/0x590 mm/slub.c:6692
            nvif_vmm_dtor+0x31/0x50 drivers/gpu/drm/nouveau/nvif/vmm.c:194
            nouveau_vmm_fini+0x16/0x50 drivers/gpu/drm/nouveau/nouveau_vmm.c:127
            nouveau_cli_fini+0x10e/0x210 drivers/gpu/drm/nouveau/nouveau_drm.c:225
            nouveau_cli_init+0x593/0xe10 drivers/gpu/drm/nouveau/nouveau_drm.c:324
            nouveau_drm_open+0x236/0x740 drivers/gpu/drm/nouveau/nouveau_drm.c:1243
            drm_file_alloc+0x5f2/0xad0 drivers/gpu/drm/drm_file.c:176
            drm_open_helper+0x1d7/0x4a0 drivers/gpu/drm/drm_file.c:335
            drm_open+0x190/0x3d0 drivers/gpu/drm/drm_file.c:388
            drm_stub_open+0x1f2/0x460 drivers/gpu/drm/drm_drv.c:1211
            chrdev_open+0x21c/0x660 fs/char_dev.c:411
            do_dentry_open+0x59d/0x12b0 fs/open.c:947
            vfs_open+0x82/0x390 fs/open.c:1052
            do_open fs/namei.c:4700 [inline]
            path_openat+0x2345/0x3420 fs/namei.c:4863
            do_file_open+0x207/0x460 fs/namei.c:4892
            do_sys_openat2+0xd1/0x1d0 fs/open.c:1368
            do_sys_open fs/open.c:1374 [inline]
            __do_sys_openat fs/open.c:1390 [inline]
            __se_sys_openat fs/open.c:1385 [inline]
            __x64_sys_openat+0x144/0x200 fs/open.c:1385
            do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
            do_syscall_64+0x115/0x690 arch/x86/entry/syscall_64.c:94
            entry_SYSCALL_64_after_hwframe+0x77/0x7f
    
            The buggy address belongs to the object at ffff888010fcdc30
            which belongs to the cache kmalloc-16 of size 16
            The buggy address is located 0 bytes inside of
            16-byte region [ffff888010fcdc30, ffff888010fcdc40)
    
            The buggy address belongs to the physical page:
            page: refcount:0 mapcount:0 mapping:0000000000000000 index:0x0 pfn:0x10fcd
            flags: 0x100000000000000(node=0|zone=1)
            page_type: f5(slab)
            raw: 0100000000000000 ffff88800d441640 dead000000000100 dead000000000122
            raw: 0000000000000000 0000000000550055 00000000f5000000 0000000000000000
            page dumped because: kasan: bad access detected
    
            Memory state around the buggy address:
            ffff888010fcdb00: fc fc 00 04 fc fc fc fc fa fb fc fc fc fc fa fb
            ffff888010fcdb80: fc fc fc fc fa fb fc fc fc fc 00 07 fc fc fc fc
            >ffff888010fcdc00: fa fb fc fc fc fc fa fb fc fc fc fc fa fb fc fc
                                                                                    ^
            ffff888010fcdc80: fc fc fa fb fc fc fc fc 00 04 fc fc fc fc fa fb
            ffff888010fcdd00: fc fc fc fc 00 00 fc fc fc fc fa fb fc fc fc fc
    
    Fix by removing the redundant teardown in nouveau_drm_open(),
    since nouveau_cli_init() already does the cleanup work.
    Also clear vmm->page after its freeing.
    
    Cc: [email protected]
    Fixes: 20d8a88e557a ("drm/nouveau: tidy up the client init/fini interfaces")
    Signed-off-by: Peiyang He <[email protected]>
    Assisted-by: LLM
    Reviewed-by: Lyude Paul <[email protected]>
    Signed-off-by: Lyude Paul <[email protected]>
    Link: https://patch.msgid.link/03BA723D9E5FF725+20260916103138.2651605-1-peiyang_he@smail.nju.edu.cn
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/nouveau: Fix gem reference leak in validate_init() [+ + +]
Author: Wentao Liang <[email protected]>
Date:   Wed Sep 16 18:02:02 2026 +0000

    drm/nouveau: Fix gem reference leak in validate_init()
    
    commit 5ea72f7b7139b123713a7983448f910bc4514d9e upstream.
    
    On the ttm_bo_reserve() failure and "vma not found" error paths, the
    loop breaks without adding the looked-up object to any validate list,
    so the reference taken by drm_gem_object_lookup() is never released;
    validate_fini() only walks the spliced lists.  Drop the reference
    before breaking out on both paths.
    
    Fixes: 19ca10d82e33bcfe ("drm/nouveau/gem: lookup VMAs for buffers referenced by pushbuf ioctl")
    Cc: [email protected]
    Signed-off-by: Wentao Liang <[email protected]>
    Reviewed-by: Lyude Paul <[email protected]>
    Signed-off-by: Lyude Paul <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/nouveau: Fix runtime PM leak in nouveau_connector_detect() [+ + +]
Author: Wentao Liang <[email protected]>
Date:   Wed Sep 16 18:03:42 2026 +0000

    drm/nouveau: Fix runtime PM leak in nouveau_connector_detect()
    
    commit 1e04611d3735543bd80a67d9d13dc13f503746fb upstream.
    
    If nvif_outp_edid_get() fails, nouveau_connector_detect() returns
    early without dropping the runtime PM reference taken at the start
    of the function, keeping the device powered on until the next
    successful detect.
    
    Balance the reference on the error path like the other exit paths
    do.
    
    Fixes: 0cd7e0718139 ("drm/nouveau/disp: add output method to fetch edid")
    Cc: [email protected]
    Signed-off-by: Wentao Liang <[email protected]>
    Reviewed-by: Lyude Paul <[email protected]>
    Signed-off-by: Lyude Paul <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/nouveau: RCU-free the scheduler-containing nouveau_sched [+ + +]
Author: Jonghyuk Kim(MalHyuk) <[email protected]>
Date:   Wed Sep 2 10:27:16 2026 +0900

    drm/nouveau: RCU-free the scheduler-containing nouveau_sched
    
    commit f7eae6d8d768fabd6b59779ca7da79e02c74e113 upstream.
    
    struct nouveau_sched embeds a struct drm_gpu_scheduler (base).
    nouveau_sched_destroy() calls nouveau_sched_fini() (which does
    drm_sched_fini(&sched->base)) and then frees the object with plain
    kfree(sched).
    
    drm_sched_fence_get_timeline_name() returns fence->sched->name, and the
    scheduler fence keeps a .release callback so it is not ops-detached on
    signalling.  A finished fence exported to userspace via drm_syncobj /
    sync_file therefore keeps pointing at &sched->base after nouveau_sched_destroy(),
    and a later get_timeline_name() -- reachable unprivileged through
    SYNC_IOC_FILE_INFO -- dereferences freed memory (KASAN slab-use-after-free
    read).
    
    Per the dma-fence lifetime contract the exporter must keep the data backing a
    signalled fence alive for an RCU grace period.  Free the scheduler-containing
    object with kfree_rcu() instead of kfree().
    
    Fixes: 5f03a507b29e ("drm/nouveau: implement 1:1 scheduler - entity relationship")
    Cc: [email protected]
    Signed-off-by: Jonghyuk Kim(MalHyuk) <[email protected]>
    Reviewed-by: Lyude Paul <[email protected]>
    Signed-off-by: Lyude Paul <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
drm/pagemap: dma-unmap pages before handling migration errors [+ + +]
Author: Matthew Brost <[email protected]>
Date:   Tue Sep 1 23:35:03 2026 -0700

    drm/pagemap: dma-unmap pages before handling migration errors
    
    [ Upstream commit 9e6372ec2a3990662ae0a67f56ac0aee19848d5b ]
    
    drm_pagemap_migrate_unmap_pages() relies on the pages array to determine
    which pages require DMA unmapping. However,
    drm_pagemap_migration_unlock_put_pages() clears the array as part of its
    cleanup, leaving drm_pagemap_migrate_unmap_pages() with no valid page
    information if it is called afterward.
    
    Call drm_pagemap_migrate_unmap_pages() before
    drm_pagemap_migration_unlock_put_pages() so the pages array remains
    valid during DMA unmapping.
    
    Reported-by: Sashiko <[email protected]>
    Fixes: f86ad0ed620c ("drm/gpusvm, drm/pagemap: Move migration functionality to drm_pagemap")
    Cc: [email protected]
    Signed-off-by: Matthew Brost <[email protected]>
    Reviewed-by: Himal Prasad Ghimiray <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

 
drm/virtio: fix memory leak of fence event on execbuffer failure [+ + +]
Author: Peiyang He <[email protected]>
Date:   Tue Sep 8 20:13:23 2026 +0800

    drm/virtio: fix memory leak of fence event on execbuffer failure
    
    commit b74aad23d99b279bb34d135795f39a6d8ecdc075 upstream.
    
    virtio_gpu_execbuffer_ioctl() reserves a DRM event with
    drm_event_reserve_init() when VIRTGPU_EXECBUF_RING_IDX selects a ring that
    userspace has enabled polling for. virtio_gpu_init_submit() does this
    before the BO handles, the command buffer, the syncobj arrays and the
    in-fence are processed, so every later error path runs with the event
    already pending, including plain argument validation failures such as an
    invalid bo_handle or an in-syncobj that carries no fence.
    
    On those paths, virtio_gpu_cleanup_submit() drops the out-fence without
    cancelling the event. The fence is freed without ever having been emitted,
    taking the only driver-side pointer to the event with it. Closing the DRM
    file does not help. drm_events_release() unlinks pending events but
    deliberately leaves the freeing to the driver's later drm_send_event(),
    which never runs for an orphaned event, so the allocation is leaked
    permanently.
    
    Found when fuzzing the virtio driver with Syzkaller:
    
            BUG: memory leak
            unreferenced object 0xffff88802c176e80 (size 96):
            comm "syz.1.367", pid 10561, jiffies 4294960122
            hex dump (first 32 bytes):
                    00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00  ................
                    c8 6e 17 2c 80 88 ff ff 00 00 00 00 00 00 00 00  .n.,............
            backtrace (crc e1973c6b):
                    kmemleak_alloc_recursive include/linux/kmemleak.h:44 [inline]
                    slab_post_alloc_hook mm/slub.c:4597 [inline]
                    slab_alloc_node mm/slub.c:4917 [inline]
                    __kmalloc_cache_noprof+0x49d/0x6f0 mm/slub.c:5485
                    _kmalloc_noprof include/linux/slab.h:988 [inline]
                    _kzalloc_noprof include/linux/slab.h:1309 [inline]
                    virtio_gpu_fence_event_create drivers/gpu/drm/virtio/virtgpu_submit.c:282 [inline]
                    virtio_gpu_init_submit drivers/gpu/drm/virtio/virtgpu_submit.c:398 [inline]
                    virtio_gpu_execbuffer_ioctl+0xbbf/0x1aa0 drivers/gpu/drm/virtio/virtgpu_submit.c:505
                    drm_ioctl_kernel+0x1f4/0x3e0 drivers/gpu/drm/drm_ioctl.c:817
                    drm_ioctl+0x5f4/0xc70 drivers/gpu/drm/drm_ioctl.c:914
                    vfs_ioctl fs/ioctl.c:51 [inline]
                    __do_sys_ioctl fs/ioctl.c:597 [inline]
                    __se_sys_ioctl fs/ioctl.c:583 [inline]
                    __x64_sys_ioctl+0x18e/0x210 fs/ioctl.c:583
                    do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
                    do_syscall_64+0x116/0x800 arch/x86/entry/syscall_64.c:94
                    entry_SYSCALL_64_after_hwframe+0x77/0x7f
    
    Fix by cancelling and freeing the DRM event on the execbuffer error path
    before dropping the fence. Clear the fence's event pointer after
    cancellation so it does not retain a dangling pointer.
    
    Fixes: cd7f5ca33585 ("drm/virtio: implement context init: add virtio_gpu_fence_event")
    Cc: [email protected]
    Signed-off-by: Peiyang He <[email protected]>
    Assisted-by: Codex:gpt-5.6-luna
    Signed-off-by: Dmitry Osipenko <[email protected]>
    Link: https://patch.msgid.link/D320EAB5680C1411+20260908121323.2405044-1-peiyang_he@smail.nju.edu.cn
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/virtio: fix NULL pointer dereference on fence allocation failure [+ + +]
Author: Peiyang He <[email protected]>
Date:   Wed Sep 9 17:11:14 2026 +0800

    drm/virtio: fix NULL pointer dereference on fence allocation failure
    
    commit 846b3c64fe3e77d9db20a7e3e62dbbb637c773e1 upstream.
    
    virtio_gpu_fence_alloc() can fail due to memory pressure and return NULL,
    but its caller like virtio_gpu_init_submit() never checks it. Later,
    virtio_gpu_init_submit() passes the NULL fence to
    virtio_gpu_fence_event_create(), which unconditionally dereferences it.
    
    Found when fuzzing the virtio driver with Syzkaller:
    
            Oops: general protection fault, probably for non-canonical address 0xdffffc0000000012: 0000 [#1] SMP KASAN NOPTI
            KASAN: null-ptr-deref in range [0x0000000000000090-0x0000000000000097]
            CPU: 1 UID: 0 PID: 9991 Comm: syz.0.121 Not tainted 7.2.0 #4 PREEMPT(full)
            Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS rel-1.17.0-0-gb52ca86e094d-prebuilt.qemu.org 04/01/2014
            RIP: 0010:virtio_gpu_fence_event_create drivers/gpu/drm/virtio/virtgpu_submit.c:295 [inline]
            RIP: 0010:virtio_gpu_init_submit drivers/gpu/drm/virtio/virtgpu_submit.c:398 [inline]
            RIP: 0010:virtio_gpu_execbuffer_ioctl+0xc78/0x1aa0 drivers/gpu/drm/virtio/virtgpu_submit.c:505
            Code: 85 ed 0f 85 21 09 00 00 e8 05 5a c9 fb 48 8b 44 24 10 48 8d b8 90 00 00 00 48 b8 00 00 00 00 00 fc ff df 48 89 fa 48 c1 ea 03 <80> 3c 02 00 0f 85 9a 0d 00 00 48 8b 44 24 10 4c 89 b0 90 00 00 00
            RSP: 0018:ffffc900039dfad0 EFLAGS: 00010216
            RAX: dffffc0000000000 RBX: ffffc900039dfdd8 RCX: ffffffff85f6fd3d
            RDX: 0000000000000012 RSI: ffffffff85f6fd4b RDI: 0000000000000090
            RBP: 0000000000000000 R0virtio_gpu_virgl_process_cmd: ctrl 0x102, error 0x1203
            R10: 0000000000000000 R11: 0000000000000000 R12: ffff8880132c4000
            R13: 0000000000000000 R14: ffff888073b6c700 R15: 000000000000003b
            FS:  00007fab480b96c0(0000) GS:ffff8880eb6e9000(0000) knlGS:0000000000000000
            CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
            CR2: 00007effbf5e55a8 CR3: 0000000048d19000 CR4: 0000000000350ef0
            Call Trace:
            <TASK>
            drm_ioctl_kernel+0x1f4/0x3e0 drivers/gpu/drm/drm_ioctl.c:817
            drm_ioctl+0x5f4/0xc70 drivers/gpu/drm/drm_ioctl.c:914
            vfs_ioctl fs/ioctl.c:51 [inline]
            __do_sys_ioctl fs/ioctl.c:597 [inline]
            __se_sys_ioctl fs/ioctl.c:583 [inline]
            __x64_sys_ioctl+0x18e/0x210 fs/ioctl.c:583
            do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
            do_syscall_64+0x116/0x800 arch/x86/entry/syscall_64.c:94
            entry_SYSCALL_64_after_hwframe+0x77/0x7f
            RIP: 0033:0x7fab471a82bd
            Code: ff c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 b0 ff ff ff f7 d8 64 89 01 48
            RSP: 002b:00007fab480b9018 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
            RAX: ffffffffffffffda RBX: 00007fab47435fa0 RCX: 00007fab471a82bd
            RDX: 00002000000000c0 RSI: 00000000c0406442 RDI: 0000000000000003
            RBP: 00007fab480b9080 R08: 0000000000000000 R09: 0000000000000000
            R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000001
            R13: 00007fab47436038 R14: 00007fab47435fa0 R15: 00007ffe85ab0740
            </TASK>
            Modules linked in:
            ---[ end trace 0000000000000000 ]---
            RIP: 0010:virtio_gpu_fence_event_create drivers/gpu/drm/virtio/virtgpu_submit.c:295 [inline]
            RIP: 0010:virtio_gpu_init_submit drivers/gpu/drm/virtio/virtgpu_submit.c:398 [inline]
            RIP: 0010:virtio_gpu_execbuffer_ioctl+0xc78/0x1aa0 drivers/gpu/drm/virtio/virtgpu_submit.c:505
            Code: 85 ed 0f 85 21 09 00 00 e8 05 5a c9 fb 48 8b 44 24 10 48 8d b8 90 00 00 00 48 b8 00 00 00 00 00 fc ff df 48 89 fa 48 c1 ea 03 <80> 3c 02 00 0f 85 9a 0d 00 00 48 8b 44 24 10 4c 89 b0 90 00 00 00
            RSP: 0018:ffffc900039dfad0 EFLAGS: 00010216
            RAX: dffffc0000000000 RBX: ffffc900039dfdd8 RCX: ffffffff85f6fd3d
            RDX: 0000000000000012 RSI: ffffffff85f6fd4b RDI: 0000000000000090
            RBP: 0000000000000000 R08: 0000000000000005 R09: 0000000000000000
            R10: 0000000000000000 R11: 0000000000000000 R12: ffff8880132c4000
            R13: 0000000000000000 R14: ffff888073b6c700 R15: 000000000000003b
            FS:  00007fab480b96c0(0000) GS:ffff888098ae9000(0000) knlGS:0000000000000000
            CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
            CR2: 00007f24c3759000 CR3: 0000000048d19000 CR4: 0000000000350ef0
            ----------------
            Code disassembly (best guess):
            0:      85 ed                   test   %ebp,%ebp
            2:      0f 85 21 09 00 00       jne    0x929
            8:      e8 05 5a c9 fb          call   0xfbc95a12
            d:      48 8b 44 24 10          mov    0x10(%rsp),%rax
            12:     48 8d b8 90 00 00 00    lea    0x90(%rax),%rdi
            19:     48 b8 00 00 00 00 00    movabs $0xdffffc0000000000,%rax
            20:     fc ff df
            23:     48 89 fa                mov    %rdi,%rdx
            26:     48 c1 ea 03             shr    $0x3,%rdx
            * 2a:   80 3c 02 00             cmpb   $0x0,(%rdx,%rax,1) <-- trapping instruction
            2e:     0f 85 9a 0d 00 00       jne    0xdce
            34:     48 8b 44 24 10          mov    0x10(%rsp),%rax
            39:     4c 89 b0 90 00 00 00    mov    %r14,0x90(%rax)
    
    Fix by checking virtio_gpu_fence_alloc() in virtio_gpu_init_submit() and
    returning -ENOMEM before any later code can dereference the NULL fence.
    
    Fixes: 70d1ace56db6 ("drm/virtio: Conditionally allocate virtio_gpu_fence")
    Cc: [email protected]
    Signed-off-by: Peiyang He <[email protected]>
    Assisted-by: Codex:gpt-5.5
    Signed-off-by: Dmitry Osipenko <[email protected]>
    Link: https://patch.msgid.link/00EFE4BA92889B14+20260909091114.2622550-1-peiyang_he@smail.nju.edu.cn
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/virtio: fix object leak in virtio_gpu_resource_create_ioctl() [+ + +]
Author: Junrui Luo <[email protected]>
Date:   Tue Sep 15 15:39:11 2026 +0800

    drm/virtio: fix object leak in virtio_gpu_resource_create_ioctl()
    
    [ Upstream commit 477bc3068fc3777b9d8ffd79e265b0dfdf2d3a6b ]
    
    virtio_gpu_resource_create_ioctl() calls drm_gem_object_release() on the
    drm_gem_handle_create() error path instead of dropping the reference it
    owns, so obj->funcs->free() never runs and the virtio_gpu_object, its
    pages and sg table, the resource id and the host-side resource are
    leaked.
    
    Use drm_gem_object_put() instead.
    
    Fixes: 62fb7a5e1096 ("virtio-gpu: add 3d/virgl support")
    Assisted-by: Claude:claude-opus-5
    Signed-off-by: Junrui Luo <[email protected]>
    Signed-off-by: Dmitry Osipenko <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

drm/virtio: fix object leak when drm_gem_handle_create() fails [+ + +]
Author: Junrui Luo <[email protected]>
Date:   Tue Sep 15 15:39:10 2026 +0800

    drm/virtio: fix object leak when drm_gem_handle_create() fails
    
    [ Upstream commit 36570ef2244cc4d7563b1f0157bc0f032498638c ]
    
    virtio_gpu_gem_create() owns the reference taken by
    virtio_gpu_object_create(). On the drm_gem_handle_create() error path it
    calls drm_gem_object_release() instead of dropping that reference.
    
    drm_gem_object_release() is the inverse of drm_gem_object_init() and does
    not touch the reference count or call obj->funcs->free(), so it is only
    correct as the last step of a destructor, as in
    virtio_gpu_cleanup_object(). Using it here leaves the bo at refcount 1
    with no remaining reference, so virtio_gpu_free_object() never runs and
    the shmem pages, sg table and virtio_gpu_object are leaked. Since
    virtio_gpu_object_create() has already set bo->created,
    VIRTIO_GPU_CMD_RESOURCE_UNREF is not queued either, leaking the host-side
    resource and the resource id.
    
    drm_gem_handle_create_tail() drops the handle reference on all of its
    internal error paths, so the caller only has to drop its own. Use
    drm_gem_object_put(), matching the success path below.
    
    Fixes: dc5698e80cf7 ("Add virtio gpu driver.")
    Reported-by: Yuhao Jiang <[email protected]>
    Assisted-by: Claude:claude-opus-5
    Signed-off-by: Junrui Luo <[email protected]>
    Signed-off-by: Dmitry Osipenko <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

drm/virtio: fix object leaks in virtio_gpu_resource_create_blob_ioctl() [+ + +]
Author: Junrui Luo <[email protected]>
Date:   Tue Sep 15 15:39:12 2026 +0800

    drm/virtio: fix object leaks in virtio_gpu_resource_create_blob_ioctl()
    
    [ Upstream commit 24b6d5c7641412c9ebef0d4c8b888d49a0e6b880 ]
    
    virtio_gpu_resource_create_blob_ioctl() calls drm_gem_object_release() on
    both the virtio_gpu_resource_assign_uuid() and drm_gem_handle_create()
    error paths instead of dropping the reference it owns, so
    obj->funcs->free() never runs and the virtio_gpu_object, the resource id
    and the host-side resource are leaked.
    
    Use drm_gem_object_put() instead.
    
    Fixes: 897b4d1acaf5 ("drm/virtio: implement blob resources: resource create blob ioctl")
    Assisted-by: Claude:claude-opus-5
    Signed-off-by: Junrui Luo <[email protected]>
    Signed-off-by: Dmitry Osipenko <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

drm/virtio: release the GEM object on virtio_gpu_vram_create() errors [+ + +]
Author: Junrui Luo <[email protected]>
Date:   Tue Sep 15 15:39:13 2026 +0800

    drm/virtio: release the GEM object on virtio_gpu_vram_create() errors
    
    [ Upstream commit 036d28db1818af2f9d80db771f5405da84d7732d ]
    
    virtio_gpu_vram_create() frees the object with a bare kfree(vram) on
    both error paths after drm_gem_private_object_init() has run, and on the
    second one after drm_gem_create_mmap_offset() has linked obj->vma_node
    into the device's VMA offset manager. The freed object stays in that
    interval tree, so a later lookup or insertion walks freed memory, and
    the dma_resv and gpuva lock are never destroyed.
    
    Call drm_gem_object_release() before kfree() on both paths.
    
    Fixes: 16845c5d5409 ("drm/virtio: implement blob resources: implement vram object")
    Assisted-by: Claude:claude-opus-5
    Signed-off-by: Junrui Luo <[email protected]>
    Signed-off-by: Dmitry Osipenko <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

drm/virtio: sync shmem backing on guest-bound transfers [+ + +]
Author: Benjamin Leggett <[email protected]>
Date:   Fri Aug 14 18:20:11 2026 -0400

    drm/virtio: sync shmem backing on guest-bound transfers
    
    [ Upstream commit 598c1c3e895590f845e04455d5580ea28ffde666 ]
    
    virtio_gpu_cmd_transfer_to_host_{2d,3d}() sync the shmem backing for the
    device before the transfer, but nothing syncs for the CPU when a transfer
    runs the other way. That breaks two ways. Where the DMA layer bounces, the
    device writes into the bounce buffer while the guest keeps reading the
    original pages. Where DMA is not coherent, the device writes memory while
    the CPU keeps stale cache lines, because nothing reaches
    arch_sync_dma_for_cpu(). Either way DRM_IOCTL_VIRTGPU_TRANSFER_FROM_HOST
    hands back stale data.
    
    Sashiko originally found this in
    https://lore.kernel.org/dri-devel/[email protected]
    but the suggestion there to fix this with dma_sync_sgtable_for_cpu()
    isn't a sufficient fix, for two reasons.
    
    - The transfer is asynchronous. virtio_gpu_cmd_transfer_from_host_3d() only
    queues the command, so a sync there would run before the device had written
    anything. It belongs on completion, and ahead of any fence signalling.
    A waiter woken by the fence would otherwise race the sync and read the
    backing pages regardless. It needs its own pass over the reclaim list
    rather than a step inside the existing one, because
    virtio_gpu_fence_event_process() also signals every earlier fence in the
    same context, so any entry in that loop may signal an earlier entry's
    fence.
    
    - The transfer is also partial, carrying an offset, a level and a box.
    Where the mapping bounces, a sync for the CPU copies the whole mapping
    back, so unless the mapping is primed first the regions the device did not
    write come back holding whatever the bounce buffer contained, discarding
    data the guest owned.
    
    So the fix: Prime the mapping before queueing, tag the vbuffer, and sync
    for the CPU on completion before the fence is signalled.
    
    A second transfer must not snapshot the mapping while an earlier one is
    still in flight, or the snapshot would predate whatever the CPU wrote once
    the earlier fence signalled and the later sync would discard it.
    
    To mitigate this, wait for outstanding fences under the reservation before
    priming.
    
    Neither sync copies anything unless the mapping genuinely bounces:
    swiotlb_sync_single_for_cpu() and its Xen counterpart look the address up
    in the bounce pool and return early when it is absent. On a platform with
    non-coherent DMA they still perform the necessary cache maintenance.
    
    The range cannot be narrowed to the box, since for a non-blob resource
    virtio_gpu_transfer_from_host_ioctl() rejects a caller-supplied stride and
    layer_stride, leaving the layout to the host and the guest with no way to
    work out which bytes the device writes. A host3d guest blob does carry
    both, so its extent could be bounded, but the sync is left whole there too
    rather than special-cased: priming makes the untouched regions round-trip
    unchanged either way.
    
    Behaviour changes worth noting:
    
    - TRANSFER_FROM_HOST can now block, where before it returned as soon as the
    command was queued. Repeated readbacks of one resource serialise, and a
    readback can wait behind an earlier queued command that touched it, since
    virtio_gpu_array_add_fence() tags uploads, execbufs and plane flushes
    alike with DMA_RESV_USAGE_WRITE. -ERESTARTSYS was already possible here
    via dma_resv_lock_interruptible().
    
    - A CPU write racing an in-flight transfer to the same resource is now
    lost, where before it survived and the transfer was lost instead. Priming
    captures the pages as of queueing, so a write landing before completion is
    overwritten by the sync.
    
    - TRANSFER_TO_HOST can also block now, but only while a guest-bound
    transfer on the same resource is outstanding, which happens only for
    callers that issue both without waiting.
    
    - Where a batch of completions contains a guest-bound transfer, the sync
    pass delays fence signalling for the whole batch. Only bounced pages are
    copied and the swiotlb pool bounds it. A batch with no such transfer is
    unaffected.
    
    Tested under QEMU on x86 with swiotlb=force and virtio-vga-gl
    iommu_platform=on, which forces both preconditions required to hit the
    original bug.
    
    Fixes: a3b815f09bb8 ("drm/virtio: add iommu support.")
    Reported-by: Sashiko AI review <[email protected]>
    Closes: https://lore.kernel.org/dri-devel/[email protected]/
    Signed-off-by: Benjamin Leggett <[email protected]>
    Signed-off-by: Dmitry Osipenko <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

 
drm/xe/gt_throttle: Report power brake as a throttle reason on CRI [+ + +]
Author: Sk Anirban <[email protected]>
Date:   Wed Sep 9 17:19:32 2026 +0530

    drm/xe/gt_throttle: Report power brake as a throttle reason on CRI
    
    [ Upstream commit ea4debcd8016f73c5dee3a29250a3d7977f015ef ]
    
    CRI defines bit 5 of the perf limit reasons register as a power brake
    (PWRBRK) indicator. Add PWRBRK_MASK and a reason_pwrbrk sysfs attribute
    for CRI in place of reason_ratl.
    
    Signed-off-by: Sk Anirban <[email protected]>
    Fixes: 8578e6d0546c ("drm/xe/gt_throttle: Drop individual show functions")
    Reviewed-by: Raag Jadav <[email protected]>
    Signed-off-by: Matthew Brost <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    (cherry picked from commit e199c851c0ab608a0ca89e7be1756a1461b41a2c)
    Signed-off-by: Rodrigo Vivi <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
drm/xe/vm: nuke PTs only after unlinking contested VMAs [+ + +]
Author: Matthew Auld <[email protected]>
Date:   Fri Sep 18 14:10:35 2026 +0100

    drm/xe/vm: nuke PTs only after unlinking contested VMAs
    
    commit 24a22fb3c731474b68e986af6804db450fb88617 upstream.
    
    In xe_vm_close_and_put(), external-BO VMAs are queued on the contested
    list for deferred destruction via xe_vma_destroy_unlocked(). However,
    xe_vm_pt_destroy() was previously invoked before processing contested
    VMAs, destroying vm->pt_root while those VMAs were still linked to their
    respective buffer objects (vm_bo->list.gpuva).
    
    If a concurrent thread evicts one of those shared buffer objects,
    xe_bo_trigger_rebind() holding only bo->resv walks the BO's VMAs and, in
    fault mode, calls xe_vm_invalidate_vma() -> xe_pt_zap_ptes(). Because
    vm->pt_root[tile->id] is already NULL, dereferencing pt->level causes a
    NULL ptr deref.
    
    Fix this by deferring xe_vm_free_scratch() and xe_vm_pt_destroy() until
    after all contested VMAs have been unlinked and destroyed.
    
    User is reporting hitting a NULL ptr deref in xe_pt_zap_ptes(), which
    could be explained by this race.
    
    Assisted-by: LLM
    Fixes: b06d47be7c83 ("drm/xe: Port Xe to GPUVA")
    Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/9290
    Signed-off-by: Matthew Auld <[email protected]>
    Cc: Thomas Hellström <[email protected]>
    Cc: Matthew Brost <[email protected]>
    Cc: <[email protected]> # v6.12+
    Reviewed-by: Thomas Hellström <[email protected]>
    Reviewed-by: Matthew Brost <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    (cherry picked from commit c2863648959489767f08892fd6e90577d2ea0b6a)
    Signed-off-by: Rodrigo Vivi <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
drm/xe: Add wa_14025941587 to xe2, xe3 and xe3p platforms [+ + +]
Author: Tangudu Tilak Tirumalesh <[email protected]>
Date:   Wed Sep 16 15:35:45 2026 +0530

    drm/xe: Add wa_14025941587 to xe2, xe3 and xe3p platforms
    
    commit cc319238e3f6668f867beb381ce93727c69b7317 upstream.
    
    Avoid programming the IDLEDLY timer to less than 5 microseconds.
    Apply wa_14025941587 to Graphics Versions 20.01 to 35.11
    and Media Versions 13.01 to 35.03
    
    v2: Use xe_rtp_match_not_sriov_vf, move to local variable
        Remove warn and other knits            - Matt R
    
    v3: Add verbose comment - Tejas
    
    v4: Restore IDLE_DLY register on engine reset.
        Add it to GUC save-restore list. -Vivek
    
    v5: Extend WA to Media Versions 13.01 to 35.03 - Vinay
    
    v6: Avoid clearing inhibit switch - Bala
        Refactor code accordingly by adding idle_reg_val.
    
    v7: Rebased with the divide-by-zero/overflow guards living in
        a separate hardening patch.
    
    v8: Preserve the Wa_16023105232 floor (DIV_ROUND_DOWN_ULL) and the
        maxcnt == 0 guard from the hardening patch. Round up
        (DIV_ROUND_UP_ULL) the Wa_14025941587 minimum conversion instead,
        so the tick-quantized delay cannot round back below 5 us.
    
    v9: Evaluate the Wa_16023105232 xe_gt_WARN_ON() against the value
        read from hardware instead of the Wa_14025941587-bumped value,
        so it no longer fires on the driver's own floor. Re-check the
        rounded-up tick value against maxcnt and floor it if tick
        quantization pushed it back to/above maxcnt, logging via
        xe_gt_dbg since this is the driver's own value, not a hardware
        anomaly.
    
    Assisted-by: GitHub_Copilot:claude-opus-4.8
    Signed-off-by: Tangudu Tilak Tirumalesh <[email protected]>
    Reviewed-by: Vinay Belgaumkar <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Matt Roper <[email protected]>
    (cherry picked from commit 9453c528fc909076468ff10df1c2e334ca5a9b00)
    Signed-off-by: Rodrigo Vivi <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>
    Signed-off-by: Thomas Hellström <[email protected]>

drm/xe: harden adjust_idledly() against divide-by-zero and overflow [+ + +]
Author: Tangudu Tilak Tirumalesh <[email protected]>
Date:   Wed Sep 16 15:35:44 2026 +0530

    drm/xe: harden adjust_idledly() against divide-by-zero and overflow
    
    commit 90f467577ebc28c5bc4ba2f0f1b83a4419fb0740 upstream.
    
    adjust_idledly() has several corner-case issues flagged during review:
    
      1. If xe_gt_clock_init() failed to recognise the crystal clock,
         gt->info.timestamp_base is 0, which makes idledly_units_ps also 0.
         The subsequent DIV_ROUND_CLOSEST(..., idledly_units_ps) is then a
         divide-by-zero and panics the kernel.
    
      2. The tick-to-ns conversions are done in u32:
            idledly * idledly_units_ps, (maxcnt - 1) * 1000
         Both overflow u32 before DIV_ROUND_CLOSEST() sees them.
    
      3. If IDLE_WAIT_TIME reads back as 0, maxcnt evaluates to 0 and
         the maxcnt - 1 clamp wraps to 0xFFFFFFFF in u32.
    
      4. The register only stores whole ticks, so the clamped ns value has
         to be converted to ticks and back. DIV_ROUND_CLOSEST() can round
         that conversion up past maxcnt:
    
           maxcnt = 640 ns, one tick = 666664 ps
    
           clamp:        maxcnt - 1        = 639 ns
           ns -> ticks:  639000 / 666664   = 0.958 -> rounds to 1 tick
           tick -> ns:    1 * 666664 / 1000 = 667 ns
    
         667 ns is programmed into RING_IDLEDLY, but 667 >= maxcnt (640),
         so xe_gt_WARN_ON() fires again on every subsequent init.
    
    Return early if timestamp_base is 0 (the unknown-crystal path).
    Do the conversions in u64 via the *_ULL() helpers so they cannot wrap.
    Clamp with a floor (DIV_ROUND_DOWN_ULL) so the programmed delay stays
    strictly below maxcnt, and guard the maxcnt == 0 case with a zero delay
    while still writing RING_IDLEDLY so INHIBIT_SWITCH_UNTIL_PREEMPTED is
    cleared.
    
    v2: Drop the redundant warn on the timestamp_base == 0 path;
        xe_gt_clock_init() already warns on an unrecognised crystal clock.
        Keep the early return to avoid the divide-by-zero. - Vinay
    
    v3: Field-mask the RING_IDLEDLY write with REG_FIELD_PREP(IDLE_DELAY, ...)
        instead of writing the raw tick count, which could clobber
        INHIBIT_SWITCH_UNTIL_PREEMPTED and reserved bits. Split the
        inhibit-switch clear from the maxcnt clamp so a set inhibit bit no
        longer forces a needless delay overwrite when the delay itself is
        already valid. Use gt_to_xe(gt) instead of gt_to_xe(hwe->gt).
    
    Fixes: d2de4410a88f ("drm/xe: Apply Wa_16023105232")
    Cc: [email protected]
    Assisted-by: GitHub_Copilot:claude-opus-4.8
    Signed-off-by: Tangudu Tilak Tirumalesh <[email protected]>
    Reviewed-by: Vinay Belgaumkar <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Matt Roper <[email protected]>
    (cherry picked from commit d864065ea25e9d12897c175de9176ce46677e176)
    Signed-off-by: Rodrigo Vivi <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/xe: Keep walking on SVM eviction failure [+ + +]
Author: Matthew Brost <[email protected]>
Date:   Thu Sep 17 13:31:58 2026 -0700

    drm/xe: Keep walking on SVM eviction failure
    
    commit be1df8badae513e01d9575398438716cfae18655 upstream.
    
    The desired behavior for SVM eviction failures, which can occur due to
    various uncontrollable races, is for TTM to continue walking the LRU
    list and look for another eviction candidate. This is expressed by
    returning -ENOSPC from the ->move() callback.
    
    Adjust the SVM eviction failure path because of races in ->move() to
    return -ENOSPC so that TTM continues searching for another buffer to
    evict.
    
    Fixes: 3ca608dc7561 ("drm/xe: Basic SVM BO eviction")
    Cc: [email protected]
    Signed-off-by: Matthew Brost <[email protected]>
    Reviewed-by: Himal Prasad Ghimiray <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Rodrigo Vivi <[email protected]>
    (cherry picked from commit 36a86c23588b8f57c9d20feb4cf5a2ab27e3baba)
    Signed-off-by: Rodrigo Vivi <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

drm/xe: Limit sg segment size to PAGE_SIZE on Xen PV [+ + +]
Author: Szymon AcedaÅ„ski <[email protected]>
Date:   Wed Sep 16 19:30:30 2026 +0200

    drm/xe: Limit sg segment size to PAGE_SIZE on Xen PV
    
    commit 141008dec73521ccf64878517460cec8b3297251 upstream.
    
    Fix display corruption on Xen PV dom0, where DMA buffers are not
    guaranteed machine-contiguous, in which case bounce buffering kicks
    in, breaking xe's memory coherency assumptions.
    
    Apply the same workaround i915 carries in i915_sg_segment_size() since
    commit 78a07fe777c4 ("drm/i915: stop abusing swiotlb_max_segment").
    
    Fixes: dd08ebf6c352 ("drm/xe: Introduce a new DRM driver for Intel GPUs")
    Reported-by: Marek Marczykowski-Górecki <[email protected]>
    Closes: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8382
    Link: https://lore.kernel.org/xen-devel/aYtznP_tT6xNPwf-@mail-itl/
    Link: https://lore.kernel.org/all/[email protected]/ # i915 counterpart
    Cc: Christoph Hellwig <[email protected]>
    Cc: Robert Beckett <[email protected]>
    Cc: [email protected] # v6.8+
    Signed-off-by: Szymon AcedaÅ„ski <[email protected]>
    Reviewed-by: Thomas Hellström <[email protected]>
    Signed-off-by: Thomas Hellström <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    (cherry picked from commit 77f704158f099b952681f207478a22d5b8218edb)
    Signed-off-by: Rodrigo Vivi <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
eth: fbnic: Avoid rounding zero ring sizes [+ + +]
Author: Björn Töpel <[email protected]>
Date:   Fri Sep 18 13:46:40 2026 +0200

    eth: fbnic: Avoid rounding zero ring sizes
    
    [ Upstream commit 0160953d8eec75c3c55562158c46442ff1fd410b ]
    
    roundup_pow_of_two() is undefined for zero. ethtool permits a zero ring
    size to reach the driver, where the minimum-size check should reject it.
    
    Leave zero unchanged while rounding nonzero ring sizes. The minimum-size
    check then rejects zero deterministically without changing the established
    behavior for other values.
    
    Fixes: 6cbf18a05c06 ("eth: fbnic: support ring size configuration")
    Reported-by: Sashiko <[email protected]>
    Link: https://lore.kernel.org/netdev/[email protected]/
    Suggested-by: Alexander Duyck <[email protected]>
    Signed-off-by: Björn Töpel <[email protected]>
    Reviewed-by: Joe Damato <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

eth: fbnic: Fix payload page pool error cleanup [+ + +]
Author: Björn Töpel <[email protected]>
Date:   Tue Sep 15 12:49:15 2026 +0200

    eth: fbnic: Fix payload page pool error cleanup
    
    [ Upstream commit 8e0b235bd918d06f54ba8fddd2c3ddc36ca59c15 ]
    
    The payload page pool pointer contains an error pointer when its
    allocation fails. The cleanup path passes that error pointer to
    page_pool_destroy() instead of destroying the header page pool. This
    can dereference the error pointer and leave the header page pool
    allocated.
    
    Destroy the header page pool instead.
    
    Fixes: 8a11010fdd96 ("eth: fbnic: allocate unreadable page pool for the payloads")
    Reported-by: Sashiko <[email protected]>
    Link: https://lore.kernel.org/netdev/[email protected]/
    Signed-off-by: Björn Töpel <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

eth: fbnic: Handle FW mailbox completions flagged with an error [+ + +]
Author: Alexander Duyck <[email protected]>
Date:   Mon Sep 14 14:10:33 2026 -0700

    eth: fbnic: Handle FW mailbox completions flagged with an error
    
    [ Upstream commit 1b97a269a5bdde20d4e69511f27649c9cb82b7c7 ]
    
    The firmware can complete a mailbox descriptor while also setting FW_ERR
    to indicate it could not process the request, for example on a mailbox
    DMA error. The completion carries no valid data.
    
    The driver did not check FW_ERR. On the Rx mailbox it would sync and
    parse the stale page as a normal message, and on the Tx mailbox it
    silently freed the request. If the initial capabilities exchange in
    fbnic_mbx_poll_tx_ready() hit FW_ERR -- on the Tx request or on the Rx
    response descriptor -- no response was parsed and the poll spun until it
    timed out even though the ring was healthy.
    
    Check FW_ERR on both mailboxes. Count it per-mailbox in
    fbnic_fw_mbx.resp_error, which is also shown in debugfs, warn (rate
    limited, since the bit is firmware controlled), and drop the Rx page
    instead of parsing it.
    
    In fbnic_mbx_poll_tx_ready() re-issue the capabilities request when
    either the Tx or the Rx resp_error counter advances, so a FW_ERR on the
    request or on its response triggers a retry rather than a timeout. A
    valid capabilities response is honored before the retry check, so a
    response parsed in the same poll as an unrelated FW_ERR is not discarded.
    The counters are mailbox-wide rather than keyed to the capabilities
    request; that is sufficient here because the exchange runs during
    bring-up before any other mailbox traffic, and any spurious retry is
    bounded by the existing 10s timeout.
    
    Fixes: da3cde08209e ("eth: fbnic: Add FW communication mechanism")
    Signed-off-by: Alexander Duyck <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/178942023343.7700.9423398932961964439.stgit@ahduyck-xeon-server.home.arpa
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

eth: fbnic: Handle maximum standalone channels [+ + +]
Author: Björn Töpel <[email protected]>
Date:   Mon Sep 14 14:10:04 2026 -0700

    eth: fbnic: Handle maximum standalone channels
    
    [ Upstream commit 1f4c73064a50f53d596c6f1d06d2d700f43c4b32 ]
    
    Standalone channels use one NAPI vector for each Tx and Rx queue.
    fbnic's allocation path excludes FBNIC_MAX_TXQS from that layout. A
    64-Tx/64-Rx configuration therefore records 128 vectors but allocates
    only 64, leaving NULL entries that resource setup dereferences.
    
    Include the maximum vector count in standalone allocation.
    
    Fixes: bc6107771bb4 ("eth: fbnic: Allocate a netdevice and napi vectors with queues")
    Signed-off-by: Björn Töpel <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/178942020457.7700.13129750616387075931.stgit@ahduyck-xeon-server.home.arpa
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

eth: fbnic: reset num_napi when the napi vectors are freed [+ + +]
Author: Alexander Duyck <[email protected]>
Date:   Mon Sep 14 14:10:18 2026 -0700

    eth: fbnic: reset num_napi when the napi vectors are freed
    
    [ Upstream commit 4bcc4a92c603fe7f062cea22e20da2e0ad6b12c3 ]
    
    fbn->num_napi is the count of live napi vectors, each of which owns an
    IRQ.  The PM path had freed them without clearing the count.
    fbnic_pm_suspend() tears the datapath down via ndo_stop() and frees the
    IRQs, but leaves netif_running() true so resume knows to re-open.  Resume
    rebuilds the datapath in __fbnic_pm_resume() and fbnic_reset_queues() sets
    num_napi and __fbnic_open() re-allocates the vectors.
    
    When the datapath is torn down but never rebuilt, num_napi is left
    pointing at freed vectors under 2 different scenarios:
     - a PCIe error recovery that fails (fbnic_err_slot_reset() ->
       __fbnic_pm_resume() returns an error -> PCI_ERS_RESULT_DISCONNECT), so
       .resume never runs; or
     - an __fbnic_open() that fails partway on resume and unwinds, freeing
       the vectors after fbnic_reset_queues() has already set num_napi.
    
    The netdev is then running with num_napi > 0 but napi[] freed, and the
    eventual remove/unbind close re-enters fbnic_down() -> fbnic_dbg_down()
    and dereferences the freed vectors:
      BUG: kernel NULL pointer dereference, address: 0000000000000210
      RIP: fbnic_dbg_down+0x28
    
    Clear num_napi when the vectors are freed: in the suspend teardown (a
    good resume re-establishes it before __fbnic_open()) and on the resume
    open failure.  A redundant ndo_stop() then walks an empty napi[].  The
    normal ndo_stop() down/up cycle is untouched and keeps num_napi for the
    next ndo_open().
    
    Fixes: bc6107771bb4 ("eth: fbnic: Allocate a netdevice and napi vectors with queues")
    Signed-off-by: Alexander Duyck <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/178942021809.7700.10804028989308077839.stgit@ahduyck-xeon-server.home.arpa
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

eth: fbnic: Set AW_FLUSH_MODE alongside AW_FLUSH when flushing the mailbox [+ + +]
Author: Alexander Duyck <[email protected]>
Date:   Mon Sep 14 14:10:25 2026 -0700

    eth: fbnic: Set AW_FLUSH_MODE alongside AW_FLUSH when flushing the mailbox
    
    [ Upstream commit 8947f13e436a4ff5eed9f8f019b2865a07af4bb2 ]
    
    When tearing down the FW mailbox Rx ring, fbnic_mbx_reset_desc_ring()
    writes AW_CFG with FLUSH set and everything else, BME included, cleared.
    Clearing BME halts the device's writes to the host but leaves the staged
    requests parked in the PUL write pipeline rather than draining them, so
    on the write path FLUSH alone never terminates the outstanding requests
    and the flush the firmware waits on never completes.
    
    Add the FLUSH_MODE definition and set both bits so the staged writes
    drain out of the pipeline on their own. BME stays cleared, so nothing
    lands on the host; it is restored later in fbnic_mbx_init_desc_ring()
    when the ring is rebuilt, once the outstanding writes are gone.
    
    The read path is unaffected. AR_CFG has no equivalent mode bit and
    AR_FLUSH terminates the outstanding reads by itself, so it is left as
    is.
    
    Both writes remain plain stores rather than read-modify-writes. That is
    deliberate: the matching write in fbnic_mbx_init_desc_ring() restores
    BME and the TLP attributes, and clears both flush bits as a side effect.
    
    Fixes: 3b12f00ddd08 ("fbnic: Gate AXI read/write enabling on FW mailbox")
    Signed-off-by: Alexander Duyck <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/178942022583.7700.11050671998277309744.stgit@ahduyck-xeon-server.home.arpa
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

eth: fbnic: use the Rx queue napi pointer to find the napi vector [+ + +]
Author: Alexander Duyck <[email protected]>
Date:   Mon Sep 14 14:10:11 2026 -0700

    eth: fbnic: use the Rx queue napi pointer to find the napi vector
    
    [ Upstream commit b5d9e9d4d0c13bc8b60d8d97e7a07fb25fea639e ]
    
    The queue management ndos pick the napi vector for an Rx queue with:
            nv = fbn->napi[idx % fbn->num_napi];
    
    The issue is this is only correct in the cases where there are no
    standalone Tx vectors. In those cases we were allocating the Tx vectors
    first and then the Rx so the queues would be pointing to Tx NAPI vectors
    instead of the Rx ones.
    
    The mapping the ndos want is already recorded.  fbnic_set_netif_napi()
    publishes it with netif_queue_set_napi(), which stores the napi pointer
    in netdev_rx_queue.napi, and fbnic_reset_netif_napi() clears it again.
    Both run under the netdev instance lock that the queue management ndos
    also hold, so the pointer can be read directly.
    
    Use it and drop the divide.  The pointer is NULL exactly while the
    datapath is down, so fbnic_queue_mem_alloc() can reject that case rather
    than reaching into freed state: netdev_rx_queue_restart() calls it
    before it tests netif_running(), and fbnic_pm_suspend() leaves
    netif_running() true across a PCIe recovery that never completes, so a
    queue restart can arrive after fbnic_stop() has freed the rings and the
    vectors.  fbnic_stop() clears the association in
    fbnic_reset_netif_queues() before fbnic_free_napi_vectors(), so the
    NULL is always published first.  fbnic_queue_start() and
    fbnic_queue_stop() need no check of their own, as
    netdev_rx_queue_reconfig() only reaches them once fbnic_queue_mem_alloc()
    has succeeded under the same instance lock.
    
    Fixes: da43127a8edc ("eth: fbnic: support queue ops / zero-copy Rx")
    Signed-off-by: Alexander Duyck <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/178942021136.7700.4391219358260544104.stgit@ahduyck-xeon-server.home.arpa
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
firewire: cdev: fix back-transition for iso_resource_auto client resource [+ + +]
Author: Takashi Sakamoto <[email protected]>
Date:   Tue Sep 22 22:26:39 2026 +0900

    firewire: cdev: fix back-transition for iso_resource_auto client resource
    
    [ Upstream commit c6b51091cafff9ce6c03c1416aa13864d17ab97c ]
    
    The todo member of iso_resource_auto structure represents the state of the
    client resource and normally transitions in the following order:
    
        ISO_RES_AUTO_ALLOC -> ISO_RES_AUTO_REALLOC -> ISO_RES_AUTO_DEALLOC
    
    However, concurrent access from the work item and the file descriptor
    release function can cause the state to transition backwards from
    ISO_RES_AUTO_DEALLOC to ISO_RES_AUTO_REALLOC.
    
    Prevent the back-transition by checking the current state before
    updating it in the work item.
    
    Fixes: fcabbf40fae5 ("firewire: core: move allocation/reallocation paths into specific branch after isoc resource management in cdev")
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Takashi Sakamoto <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
fou: reject omitted FOU_ATTR_IPPROTO on FOU_ENCAP_DIRECT [+ + +]
Author: Hui Peng <[email protected]>
Date:   Mon Sep 21 04:59:20 2026 +0000

    fou: reject omitted FOU_ATTR_IPPROTO on FOU_ENCAP_DIRECT
    
    commit d22609f3d13fc5baacd92c222731b03c593401db upstream.
    
    Commit 7a9bc9e3f423 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.") added
    NLA_POLICY_MIN(NLA_U8, 1) to fou_nl_policy[FOU_ATTR_IPPROTO], which
    rejects an explicitly supplied FOU_ATTR_IPPROTO == 0 attribute with
    -ERANGE.
    
    However, FOU_ATTR_IPPROTO is an optional netlink attribute. When a user
    sends FOU_CMD_ADD with FOU_ATTR_TYPE set to FOU_ENCAP_DIRECT and omits
    FOU_ATTR_IPPROTO entirely, nla_policy validation succeeds and
    parse_nl_config() leaves cfg->protocol as 0 (from memset(cfg, 0,
    sizeof(*cfg))). fou_create() then creates a FOU_ENCAP_DIRECT socket with
    fou->protocol == 0.
    
    In fou_udp_recv(), returning -fou->protocol to udp_queue_rcv_one_skb()
    triggers IP protocol resubmission when fou->protocol > 0, whereas
    returning 0 tells the UDP tunnel layer that the skb was consumed without
    freeing it. When fou->protocol == 0, every packet received on the socket
    returns 0 from fou_udp_recv() and leaks the sk_buff.
    
    Reject FOU_ENCAP_DIRECT when !cfg->protocol in fou_create() so that
    creating a direct encapsulation port without FOU_ATTR_IPPROTO fails with
    -EINVAL while leaving FOU_CMD_DEL and FOU_CMD_GET (which share
    parse_nl_config()) unaffected.
    
    Tested in QEMU against Linux 7.3.0-rc3 by sending a FOU_CMD_ADD Generic
    Netlink request with FOU_ATTR_PORT = 5555 and FOU_ATTR_TYPE =
    FOU_ENCAP_DIRECT while omitting FOU_ATTR_IPPROTO. On the unfixed kernel,
    FOU_CMD_ADD succeeds (err = 0), FOU_CMD_GET reports fou->type = 1 and
    fou->protocol = 0, and sending 4000 UDP packets to 127.0.0.1:5555 leaks
    all 4000 sk_buffs (SUnreclaim in /proc/meminfo grows from 41456 kB to
    59008 kB, +17552 kB); with this patch applied, FOU_CMD_ADD is rejected
    with -EINVAL (-22).
    
    Fixes: 23461551c006 ("fou: Support for foo-over-udp RX path")
    Fixes: 7a9bc9e3f423 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.")
    Cc: [email protected]
    Signed-off-by: Hui Peng <[email protected]>
    Reviewed-by: Hangbin Liu <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
fprobe: Terminate the fgraph_data list when the reservation is not filled [+ + +]
Author: David Carlier <[email protected]>
Date:   Thu Sep 17 22:24:07 2026 +0100

    fprobe: Terminate the fgraph_data list when the reservation is not filled
    
    commit 1d653a183973f5283a3db5a38cd5e195eb152244 upstream.
    
    fprobe_fgraph_entry() reserves shadow stack space for every fprobe with
    an exit handler, but only fills it for those whose entry handler returns
    0. fgraph_reserve_data() does not clear the area, so fprobe_return()
    parses the unused tail as headers left over from an earlier call, and an
    exit handler can run twice or despite its entry handler asking to skip
    it.
    
    Write a zero word after the last entry to terminate the walk. A zeroed
    slot does not decode to a NULL fprobe on the arches that encode the
    header into one unsigned long, since arch_decode_fprobe_header_fp() ORs
    in FPROBE_HEADER_MSB_PATTERN, so make read_fprobe_header() return NULL
    for a zeroed slot.
    
    Link: https://lore.kernel.org/all/[email protected]/
    
    Fixes: e0a384434ae1 ("tracing: fprobe: do not zero out unused fgraph_data")
    Cc: [email protected]
    Suggested-by: Masami Hiramatsu (Google) <[email protected]>
    Signed-off-by: David Carlier <[email protected]>
    Signed-off-by: Masami Hiramatsu (Google) <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
fs/ntfs3: use d_instantiate_new() in ntfs_create_inode() and murder syzbot's "WARNING in do_new_mount" saga [+ + +]
Author: Christian Brauner <[email protected]>
Date:   Wed Sep 9 11:03:18 2026 +0200

    fs/ntfs3: use d_instantiate_new() in ntfs_create_inode() and murder syzbot's "WARNING in do_new_mount" saga
    
    commit 1abd643f3783ea8f8e273c18697ff0413aa92dc7 upstream.
    
    ntfs_create_inode() creates a new inode via ntfs_new_inode(). It hashes
    it with insert_inode_locked() and so it's marked as I_NEW until
    unlock_new_inode().
    
    ntfs 3 calls d_instantiate() in between though... Since the dentry was
    already hashed by the lookup before the create any path walk finds it
    without touching the parent's i_rwsem and so can lock the inode.
    
    If the inode is a directory unlock_new_inode() calls
    lockdep_annotate_inode_mutex_key() and marks i_rwsem with the
    i_mutex_dir_key class.
    
    That resets the count and the owner of a lock somebody else may already
    hold by now...
    
    syzbot has been spamming us with the same godforsaken bug
    
      "WARNING in do_new_mount"
    
    since 2023. I can't take it anymore so I went looking. Afaict, syzbot's
    executor chdirs into a freshly mounted ntfs3 image, creates a
    directory and then mounts some pseudofs on it. Everytime the mkdir()
    takes longer than syzbot waits mount() runs concurrently:
    
      mkdir("./sys")                        mount(NULL, "./sys", "sysfs")
      ntfs_create_inode()
        d_instantiate()
                                            user_path_at() finds the dentry
                                            do_lock_mount()
                                              inode_lock(inode)
                                              namespace_lock()
        unlock_new_inode()
          lockdep_annotate_inode_mutex_key()
            init_rwsem(&inode->i_rwsem)
                                            unlock_mount()
                                              inode_unlock(inode)
    
    The mount side then releases a lock that according to the rwsem nobody
    holds:
    
      DEBUG_RWSEMS_WARN_ON((rwsem_owner(sem) != current) && ...):
      count = 0x0, magic = 0xffff888043a854e8, owner = 0x0,
      curr 0xffff888000244880, list empty
      WARNING: CPU: 0 PID: 5346 at kernel/locking/rwsem.c:1368 __up_write
      Call Trace:
       inode_unlock include/linux/fs.h:877 [inline]
       unlock_mount fs/namespace.c:2892 [inline]
       do_new_mount_fc fs/namespace.c:3828 [inline]
       do_new_mount+0x777/0xa40 fs/namespace.c:3887
    
    On PREEMPT_RT the same thing shows up as
    
      DEBUG_LOCKS_WARN_ON(rt_mutex_owner(lock) != current)
      WARNING: kernel/locking/rtmutex_common.h:193 at rt_mutex_slowunlock
    
    The up_write() underflows the reset count. A following inode_lock() on
    that directory then never returns. A path walk into the new directory
    racing with the mkdir() corrupts the lock the same way via
    inode_lock_shared() in lookup_slow().
    
    Switch to d_instantiate_new() and drop the trailing unlock_new_inode().
    All error paths bail out before that point with I_NEW still set and
    keep using discard_new_inode().
    
    May we never see this fscking bug report again.
    
    Link: https://patch.msgid.link/20260909-work-ntfs3-d_instantiate_new-v1-1-2db697162ce8@kernel.org
    Fixes: 82cae269cfa9 ("fs/ntfs3: Add initialization of super block")
    Reviewed-by: Jan Kara <[email protected]>
    Cc: [email protected] # v5.15+
    Reported-by: [email protected]
    Closes: https://lore.kernel.org/[email protected]
    Signed-off-by: Christian Brauner (Amutable) <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
fs: avoid repeated scans in evict_inodes() [+ + +]
Author: Julian Sun <[email protected]>
Date:   Tue Sep 15 12:49:12 2026 +0800

    fs: avoid repeated scans in evict_inodes()
    
    [ Upstream commit 7459c021874246c196f397686e100702f059b9e7 ]
    
    We observed hung tasks when users attempted to unmount a filesystem
    after its disk had been removed while still in use. During device
    removal, fs_bdev_mark_dead() calls evict_inodes() while holding s_umount.
    
    Each time evict_inodes() drops s_inode_list_lock to reschedule, it
    restarts the walk from the head of s_inodes. With many referenced inodes
    at the head of the list, these restarts repeatedly scan the same inodes
    without reclaiming them. This can keep s_umount held for a long time,
    blocking concurrent umount attempts and triggering hung-task reports.
    
    Keep the current inode, already marked I_FREEING, out of the disposal
    batch until s_inode_list_lock is reacquired. Resume the walk from this
    inode and dispose of it in a later batch or at the end of the walk.
    
    The zero-refcount and state checks under i_lock allow this walker to
    claim the inode by setting I_FREEING and removing it from the LRU.
    Other reclaimers skip the inode, leaving this walker responsible for
    eviction. Only evict() removes it from s_inodes, so keeping it out of
    the disposal batch ensures that it remains on the list while the lock
    is dropped. After reacquiring the lock, reading its current next pointer
    accounts for concurrent removal of following inodes.
    
    The existing inode lifetime rules prohibit acquiring a reference to an
    inode marked I_FREEING or I_WILL_FREE. __iget() requires its caller to
    hold i_lock and establish that taking a reference is valid. Inode lookup
    and igrab() check these flags under i_lock when acquiring a reference
    from zero. ihold() requires an existing reference, which would keep
    i_count nonzero and prevent this walker from claiming the inode. These
    rules already allow iput_final() and the inode shrinker to release
    i_lock after setting I_FREEING and before eviction completes.
    
    A temporary __iget() reference would also keep the inode on the list,
    but its release must preserve last-reference handling. Another user can
    acquire a reference, update lazy timestamps and drop its reference while
    the pin is held. If the pin becomes the last reference, dropping it with
    atomic_dec_and_test() and evicting directly bypasses iput()'s lazytime
    handling and can lose those timestamp updates.
    
    Releasing the pin with iput() preserves that handling, but does not
    guarantee eviction. fs_bdev_mark_dead() runs with SB_ACTIVE set, so iput()
    may retain the inode in cache, whereas evict_inodes() must evict eligible
    zero-reference inodes. The inode may also have been freed when iput()
    returns, so the walker cannot then use it to force eviction. Using
    I_FREEING preserves the existing eviction behavior without introducing
    an additional last-reference transition.
    
    The xfstests auto group passed on ext4 and XFS with known unrelated
    failures excluded. No new issues were observed, and the previously
    reproducible hung task no longer occurs with this patch.
    
    Fixes: ac05fbb40062 ("inode: don't softlockup when evicting inodes")
    Signed-off-by: Julian Sun <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Reviewed-by: Jan Kara <[email protected]>
    Signed-off-by: Christian Brauner (Amutable) <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
fsl/fman: Fix clk reference leak in read_dts_node() [+ + +]
Author: Wentao Liang <[email protected]>
Date:   Thu Sep 17 11:01:35 2026 +0000

    fsl/fman: Fix clk reference leak in read_dts_node()
    
    commit a644f09b2090ad22a13fbcf9d141084f573108ef upstream.
    
    of_clk_get() returns a clock with its reference count incremented, but
    read_dts_node() only uses it to read the rate and never calls clk_put().
    The clock is not stored anywhere, so the reference cannot be released
    later either.
    
    Release the clock once its rate has been read, which also covers the
    error path taken when the rate is zero.
    
    Fixes: 414fd46e7762 ("fsl/fman: Add FMan support")
    Cc: [email protected]
    Signed-off-by: Wentao Liang <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
genetlink: report the real command id for dump-only ops in policy dumps [+ + +]
Author: Jakub Kicinski <[email protected]>
Date:   Fri Sep 18 15:29:48 2026 -0700

    genetlink: report the real command id for dump-only ops in policy dumps
    
    [ Upstream commit 261e8a37ecbaf462cdf9c336d2b2f5056088401a ]
    
    The op-to-policy map a CTRL_CMD_GETPOLICY dump returns is the only way
    for userspace to find out which policy index belongs to which command.
    ctrl_dumppolicy_put_op() tags the nest with doit->cmd, but an op which
    only has a dumpit has no doit and every path which fills the split ops
    in zeroes it out, so those entries all claim to be command 0.  nlctrl's
    own CTRL_CMD_GETPOLICY and NETDEV_CMD_QSTATS_GET are both in that group:
    
      [{'family-id': 16, 'op-policy': {'do': 0, 'dump': 0, 'op-id': 3}},
       {'family-id': 16, 'op-policy': {'dump': 1, 'op-id': 0}},
    
    ctrl_fill_info() gets this right - it uses the iterator's cmd for
    CTRL_ATTR_OP_ID - so the two introspection interfaces of the same family
    contradict each other today.
    
    Pass the command in rather than reconstructing it from
    doit->cmd | dumpit->cmd inside the helper, both callers already have it.
    
    Fixes: 26588edbef60 ("genetlink: support split policies in ctrl_dumppolicy_put_op()")
    Signed-off-by: Jakub Kicinski <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
gpio: arizona: Fix runtime PM leak in arizona_gpio_direction_out() [+ + +]
Author: Wentao Liang <[email protected]>
Date:   Wed Sep 16 09:47:01 2026 +0000

    gpio: arizona: Fix runtime PM leak in arizona_gpio_direction_out()
    
    commit e9d810279f84b30738f7790c0ed15f8dd5b9024a upstream.
    
    Switching a persistent GPIO line from input to output acquires a
    runtime PM reference on the parent device, but if the subsequent
    regmap_update_bits() fails the reference is never dropped and no later
    direction_in() can balance it since the direction was never changed.
    Drop the reference on the update failure path.
    
    Fixes: 27a49ed17e22 ("gpio: arizona: Add support for GPIOs that need to be maintained")
    Cc: [email protected]
    Signed-off-by: Wentao Liang <[email protected]>
    Reviewed-by: Charles Keepax <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Bartosz Golaszewski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

gpio: cdev: fix kernel stack leak to user-space in error path [+ + +]
Author: Bartosz Golaszewski <[email protected]>
Date:   Tue Sep 22 10:52:28 2026 +0200

    gpio: cdev: fix kernel stack leak to user-space in error path
    
    commit 1feb5d39b05afd902ed9fc902ec5b15be03a4bdb upstream.
    
    If we fail to acquire the GPIO chip guard in gpio_desc_to_lineinfo(), we
    return immediately before zeroing the info struct we'll end up passing
    to the user-space later in lineinfo_get_v1(). This may leak the kernel
    stack contents. Make gpio_desc_to_lineinfo() return int so that the
    -ENODEV returned on failure to acquire the guard can be propagated to
    the callers.
    
    While not strictly necessary: move the memset() before trying to acquire
    the SRCU read lock too for good measure.
    
    Fixes: d83cee3d2bb1 ("gpio: protect the pointer to gpio_chip in gpio_device with SRCU")
    Cc: [email protected]
    Reported-by: Sashiko <[email protected]>
    Closes: https://sashiko.dev/#/patchset/20260912123529.7951-1-tzungbi%40kernel.org?part=3
    Reviewed-by: Kent Gibson <[email protected]>
    Link: https://patch.msgid.link/20260922-gpio-cdev-stack-leak-fixes-v3-1-7a0c7a4299d5@oss.qualcomm.com
    Signed-off-by: Bartosz Golaszewski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

gpio: tps65219: Fix GPIO input value reads [+ + +]
Author: Karl Mehltretter <[email protected]>
Date:   Sat Sep 19 19:10:58 2026 +0200

    gpio: tps65219: Fix GPIO input value reads
    
    commit 4cbe530c0233c7413aaaeb029a4f32dd6aadacbb upstream.
    
    TPS65219_MFP_GPIO_STATUS_MASK is already BIT(4). Passing it to BIT()
    again tests bit 16, which cannot be set in the 8-bit MFP_CTRL register,
    so GPIO0 is always reported low when configured as an input.
    
    Test the register value with the mask directly.
    
    Fixes: 57e30e00bd5b ("gpio: tps65219: add GPIO support for TPS65219 PMIC")
    Cc: [email protected]
    Assisted-by: LLM
    Signed-off-by: Karl Mehltretter <[email protected]>
    Reviewed-by: Jonathan Cormier <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Bartosz Golaszewski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

gpio: tps65219: Fix TPS65214 GPIO direction programming [+ + +]
Author: Karl Mehltretter <[email protected]>
Date:   Sat Sep 19 19:11:00 2026 +0200

    gpio: tps65219: Fix TPS65214 GPIO direction programming
    
    commit 270437f3fe62516f16482742a7762a075e7a9457 upstream.
    
    GPIO_LINE_DIRECTION_OUT and GPIO_LINE_DIRECTION_IN have the values 0
    and 1, respectively, while the TPS65214 GPIO_CONFIG field is BIT(1).
    regmap_update_bits() masks the supplied value, so passing either
    direction value clears the field and selects input mode.
    
    Translate the GPIO direction to the register encoding used by
    tps65214_gpio_get_direction(), setting GPIO_CONFIG for output and
    clearing it for input.
    
    Fixes: 1b6ab07c0c80 ("gpio: tps65219: Add support for TI TPS65214 PMIC")
    Cc: [email protected]
    Assisted-by: LLM
    Signed-off-by: Karl Mehltretter <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Bartosz Golaszewski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

gpio: tps65219: Use the variant-specific direction callback [+ + +]
Author: Karl Mehltretter <[email protected]>
Date:   Sat Sep 19 19:10:59 2026 +0200

    gpio: tps65219: Use the variant-specific direction callback
    
    commit 93cf8cedeaaa05714f709b539cfb976e0b80c830 upstream.
    
    The TPS65214 template installs its own get_direction callback because
    its direction bit is in GENERAL_CONFIG. The shared get and direction
    callbacks nevertheless call tps65219_gpio_get_direction() directly and
    interpret the unrelated TPS65219 MFP bit.
    
    On TPS65214 this can reject reads from an input and skip the change from
    input to output. Call the callback selected by the gpio_chip template
    instead.
    
    Fixes: 1b6ab07c0c80 ("gpio: tps65219: Add support for TI TPS65214 PMIC")
    Cc: [email protected]
    Assisted-by: LLM
    Signed-off-by: Karl Mehltretter <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Bartosz Golaszewski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

gpio: zynq: fix runtime PM leak on request error path [+ + +]
Author: Ridham Khurana <[email protected]>
Date:   Tue Sep 22 09:20:58 2026 +0000

    gpio: zynq: fix runtime PM leak on request error path
    
    commit e9438ab5328a177c9c0e5df87eb92a7162841e98 upstream.
    
    pm_runtime_get_sync() leaves the usage counter incremented even when it
    fails, and zynq_gpio_request() returns the error without dropping it.
    
    gpiolib does not call ->free() when ->request() fails, so zynq_gpio_free(),
    which holds the only matching pm_runtime_put(), never runs. The reference
    is leaked and the controller can no longer runtime-suspend, so its clock
    stays enabled.
    
    Switch to pm_runtime_resume_and_get(), which only increments the usage
    counter on success.
    
    Fixes: 3242ba117e9b ("gpio: Add driver for Zynq GPIO controller")
    Cc: [email protected]
    Signed-off-by: Ridham Khurana <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Bartosz Golaszewski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
gpiolib: use of_node_name if line-name is missing [+ + +]
Author: Frank Wunderlich <[email protected]>
Date:   Thu Sep 17 17:37:11 2026 +0200

    gpiolib: use of_node_name if line-name is missing
    
    [ Upstream commit d54a489c8c4b627775445d54a6cbd32f209bda8f ]
    
    Until v7.0, GPIO hogs inherited the DT node name when no line-name
    property was specified. This was implemented as a fallback in
    of_parse_own_gpio().
    
    Commit d1d564ec4992 ("gpio: move hogs into GPIO core") moved hog parsing
    into the GPIO core and removed this fallback.
    
    Consequently, GPIO hogs without a line-name property are now displayed
    with a ? in /sys/kernel/debug/gpio. Restore the old fallback.
    
    Fixes: d1d564ec4992 ("gpio: move hogs into GPIO core")
    Signed-off-by: Frank Wunderlich <[email protected]>
    Reviewed-by: Andy Shevchenko <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Bartosz Golaszewski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO [+ + +]
Author: Eric Dumazet <[email protected]>
Date:   Wed Sep 23 14:59:42 2026 +0000

    gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO
    
    [ Upstream commit 83769c23fb1879edc916a526ba424285033baf2d ]
    
    gve_can_send_tso() computes how many buffers each segment of a GSO
    packet would span, and for this it needs the length of the headers
    that the device replicates in front of every segment.
    
    It unconditionally uses skb_tcp_all_headers(), which reads the doff
    field of the TCP header. SKB_GSO_UDP_L4 packets have no TCP header:
    tcp_hdrlen() then reads one byte of the UDP payload, and header_len
    can be anything in [0, 60] instead of the transport offset plus the
    eight bytes of the UDP header that gve_prep_tso() programs into the
    TSO context descriptor.
    
    A wrong header length shifts all the segment boundaries computed in
    the loop, so the number of buffers per segment can be over or under
    estimated. In the first case, GSO is needlessly disabled for this
    packet by gve_features_check_dqo() and the stack has to segment it.
    In the second case, the driver hands the device a packet whose
    segments span more than GVE_TX_MAX_DATA_DESCS buffers.
    
    Use the UDP header length for SKB_GSO_UDP_L4 packets, matching what
    gve_prep_tso() does.
    
    Fixes: 014c607f86ab ("gve: add support for UDP GSO for DQO format")
    Closes: https://lore.kernel.org/netdev/CANn89i+MS4L60sFQ49=-f-mibeveUfcrpVkD5X+Qy6SOnEpd6w@mail.gmail.com/
    Signed-off-by: Eric Dumazet <[email protected]>
    Cc: Ankit Garg <[email protected]>
    Cc: Harshitha Ramamurthy <[email protected]>
    Cc: Joshua Washington <[email protected]>
    Cc: Willem de Bruijn <[email protected]>
    Reviewed-by: Ankit Garg <[email protected]>
    Reviewed-by: Harshitha Ramamurthy <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

gve: DQO: reject TSO packets with an out of range MSS [+ + +]
Author: Eric Dumazet <[email protected]>
Date:   Thu Sep 24 00:42:52 2026 +0000

    gve: DQO: reject TSO packets with an out of range MSS
    
    [ Upstream commit 296c83b5ccc808c080865eb20fd7a477b0355bb7 ]
    
    gve_prep_tso() notes that the device requires the MSS to be <= 9728,
    but does not enforce it, assuming the 9K MTU enforced by the hypervisor
    and the 64KB limit on TSO sizes are enough.
    
    This does not hold for packets that were not generated locally.
    A guest behind a tap, or any packet socket user, can provide an
    arbitrary gso_size in virtio_net_hdr. Layer 2 forwarding does not check
    the MTU for GSO packets (is_skb_forwardable()), and gso_features_check()
    only bounds skb->len and gso_segs, never gso_size.
    
    Such a packet reaches gve_tx_fill_tso_ctx_desc(), which puts gso_size
    into the mss field of the TSO context descriptor. This field is 14 bits
    wide, so a gso_size of 16384 is silently turned into an MSS of zero.
    
    Drop these packets from gve_prep_tso(), and make sure that
    gve_features_check_dqo() leaves their GSO bits alone: skb_segment()
    splits at gso_size regardless of the MTU, so falling back to software
    segmentation would give the device non TSO packets bigger than the
    9728 bytes it supports.
    
    Note that the device can still be given oversized non TSO packets when
    the stack segments in software for other reasons, for instance after
    TSO has been disabled with ethtool. This is a generic issue, because
    the MTU check is skipped for GSO packets in the forwarding path, and
    is addressed separately.
    
    Fixes: a57e5de476be ("gve: DQO: Add TX path")
    Signed-off-by: Eric Dumazet <[email protected]>
    Reviewed-by: Harshitha Ramamurthy <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

gve: fix TX drop when GSO MSS is too small for hw [+ + +]
Author: Eddie Phillips <[email protected]>
Date:   Thu Sep 24 00:42:51 2026 +0000

    gve: fix TX drop when GSO MSS is too small for hw
    
    [ Upstream commit 3b430ea6234087957b0d3cd181e3116722b59819 ]
    
    The device has a strict requirement that the minimum MSS
    (gso_size) for TSO/GSO packets must be at least 88 bytes. If a packet
    below this threshold is pushed to the hardware, it can cause
    hardware to silently drop the packet, leading to increased latency
    and retransmissions.
    
    Currently, this is validated too late in the transmit pipeline
    (gve_prep_tso), leading to silent drops.
    
    Fix this by moving the validation into the .ndo_features_check
    callback (gve_features_check_dqo). If we detect a GSO packet with
    a gso_size smaller than GVE_TX_MIN_TSO_MSS_DQO, we clear the GSO
    feature flags for this packet.
    
    Fixes: a57e5de476be ("gve: DQO: Add TX path")
    Signed-off-by: Eddie Phillips <[email protected]>
    Signed-off-by: Eric Dumazet <[email protected]>
    Reviewed-by: Harshitha Ramamurthy <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
HID: alps: fix use-after-free on input2 registration failure [+ + +]
Author: Chen Changcheng <[email protected]>
Date:   Fri Aug 14 15:06:21 2026 +0800

    HID: alps: fix use-after-free on input2 registration failure
    
    commit d3aba3442798ce4a4c8ce3104d7b286d61e605f9 upstream.
    
    alps_input_configured() stores data->input2 before calling
    input_register_device().  If registration fails, input_free_device()
    frees the input device but data->input2 still points to the freed memory.
    alps_input_configured() calls hid_hw_open() before allocating input2, so
    URBs are already active and raw_event can fire during the failure window.
    A U1_SP_ABSOLUTE_REPORT_ID report arriving then causes u1_raw_event()
    to dereference the freed data->input2 -> use-after-free.
    
    Fix by only storing input2 into drvdata after successful registration
    and adding a NULL guard in the raw_event path.
    
    Fixes: 2562756dde55 ("HID: add Alps I2C HID Touchpad-Stick support")
    Cc: [email protected]
    Signed-off-by: Chen Changcheng <[email protected]>
    Signed-off-by: Jiri Kosina <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

HID: alps: unregister DualPoint Stick input device on remove [+ + +]
Author: Chen Changcheng <[email protected]>
Date:   Fri Aug 14 15:06:20 2026 +0800

    HID: alps: unregister DualPoint Stick input device on remove
    
    commit aa9dde93e05a837645fdfd577eea71f0733f0694 upstream.
    
    alps_input_configured() allocates a second input device ("DualPoint
    Stick") with input_allocate_device() and registers it, but the
    alps_driver struct has no .remove handler and input2 is not tracked in
    hdev->inputs.  The default remove path (hid_hw_stop -> hidinput_disconnect)
    only iterates hdev->inputs, so input2 is never unregistered and leaks
    on every device removal.
    
    Add a .remove handler that stops the device first (preventing URB
    callbacks from touching input2 during teardown) and then unregisters
    input2.
    
    Fixes: 2562756dde55 ("HID: add Alps I2C HID Touchpad-Stick support")
    Cc: [email protected]
    Signed-off-by: Chen Changcheng <[email protected]>
    Signed-off-by: Jiri Kosina <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

HID: amd_sfh: Validate PCI BAR size before mapping [+ + +]
Author: Slawomir Stepien <[email protected]>
Date:   Mon Sep 14 11:50:07 2026 +0200

    HID: amd_sfh: Validate PCI BAR size before mapping
    
    [ Upstream commit 65bcc5f89704efe5b9d69d4ea2c1002d90c31382 ]
    
    The amd_sfh driver maps PCI BAR 2 using pcim_iomap_regions() and
    subsequently accesses MMIO registers at offsets up to 0x10958 (e.g.,
    AMD_P2C_MSG3 at 0x1068C). However, the driver never validates that the BAR
    size is large enough to cover these accesses. If the driver is bound to a
    device with a smaller BAR 2, this leads to an out-of-bounds memory access
    and a page fault during the probe function.
    
    For example, a page fault can occur when reading from privdata->mmio +
    AMD_P2C_MSG3 in mp2_select_ops():
    
      BUG: unable to handle page fault for address: ffffc9000390368c
      PGD 100000067 P4D 100000067 PUD 1012c1067 PMD 105b64067 PTE 0
      Oops: Oops: 0000 [#1] SMP KASAN NOPTI
      RIP: 0010:readl arch/x86/include/asm/io.h:59 [inline]
      RIP: 0010:mp2_select_ops drivers/hid/amd-sfh-hid/amd_sfh_pcie.c:282
      [inline]
      RIP: 0010:amd_mp2_pci_probe+0x337/0x5f0
      drivers/hid/amd-sfh-hid/amd_sfh_pcie.c:487
      Call Trace:
       <TASK>
       local_pci_probe drivers/pci/pci-driver.c:332 [inline]
       pci_call_probe drivers/pci/pci-driver.c:394 [inline]
       __pci_device_probe drivers/pci/pci-driver.c:455 [inline]
       pci_device_probe+0x431/0xc90 drivers/pci/pci-driver.c:489
    
    Fix this by verifying that the length of BAR 2 is at least 128KB before
    attempting to map it. Since the maximum accessed offset is 0x10958, and PCI
    BAR sizes are powers of 2, any legitimate hardware will have a BAR size of
    at least 128KB.
    
    Fixes: 4f567b9f8141 ("SFH: PCIe driver to add support of AMD sensor fusion hub")
    Assisted-by: Gemini:gemini-3.7-flash Gemini:gemini-3.1-pro-preview syzbot
    Reported-by: [email protected]
    Closes: https://syzkaller.appspot.com/bug?extid=4eadd4dfe9e66522bae8
    Link: https://syzkaller.appspot.com/ai_job?id=3bc1c45c-548f-4ab5-8243-d2c8ec321d6c
    Signed-off-by: Slawomir Stepien <[email protected]>
    Acked-by: Basavaraj Natikar <[email protected]>
    Link: https://syzkaller.appspot.com/bug?extid=4eadd4dfe9e66522bae8
    Signed-off-by: Jiri Kosina <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

HID: bpf: fix __hid_bpf_hw_check_params report length [+ + +]
Author: Benjamin Tissoires <[email protected]>
Date:   Fri Sep 4 14:53:00 2026 +0200

    HID: bpf: fix __hid_bpf_hw_check_params report length
    
    [ Upstream commit c4afa4862b878d56e0cc1021298794ac1b45bc49 ]
    
    Turns out that USB, I2C and other transport drivers (except uhid which
    just passes the data) still need to have the report ID in the first
    byte.
    
    Because they expect the first byte to be the report ID or 0, when the
    report ID is 0, they strip that first byte before forwarding to the
    device. This means that the transport layer forwards a buffer of size
    N-1 to the device, which gets rejected.
    
    Fixes: 5599f8019661 ("HID: bpf: export hid_hw_output_report as a BPF kfunc")
    Signed-off-by: Benjamin Tissoires <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

HID: elecom: fix bus type for M-XGL20DLBK [+ + +]
Author: Oscar Priego Verdugo <[email protected]>
Date:   Mon Aug 17 06:05:46 2026 -0600

    HID: elecom: fix bus type for M-XGL20DLBK
    
    [ Upstream commit 8e2a4b458ad25e13422bb059758c30a6562aa9cf ]
    
    The M-XGL20DLBK is matched as a USB device by hid-elecom, but
    its entry in hid_have_special_driver[] uses HID_BLUETOOTH_DEVICE.
    
    This prevents the special-driver quirk entry from matching the USB
    device handled by hid-elecom. Use HID_USB_DEVICE there as well.
    
    Fixes: 55633e681afb ("HID: elecom: add support for EX-G M-XGL20DLBK wireless mouse")
    Signed-off-by: Oscar Priego Verdugo <[email protected]>
    Signed-off-by: Jiri Kosina <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

HID: hid-oxp: use cancel_delayed_work_sync() in remove [+ + +]
Author: Tristan Madani <[email protected]>
Date:   Fri Sep 4 10:58:00 2026 +0000

    HID: hid-oxp: use cancel_delayed_work_sync() in remove
    
    commit abd24922c2a9797d6184be0561dd85e7bcdfd091 upstream.
    
    oxp_hid_remove() uses cancel_delayed_work() for all three delayed work
    items.  cancel_delayed_work() only dequeues a pending work item without
    waiting for a currently executing callback to finish.  If any of the
    work callbacks (oxp_rgb_queue_fn, oxp_btn_queue_fn, oxp_mcu_init_fn) is
    running at the time of removal, the callback continues executing
    concurrently with hid_hw_close() and hid_hw_stop(), accessing the HID
    device after it has been closed and stopped.
    
    Use cancel_delayed_work_sync() instead to ensure that any in-progress
    work callback completes before device teardown proceeds.
    
    Fixes: 84910c459d65 ("HID: hid-oxp: Add OneXPlayer configuration driver")
    Cc: [email protected]
    Signed-off-by: Tristan Madani <[email protected]>
    Reviewed-by: Derek J. Clark <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Jiri Kosina <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

HID: multitouch: Add report ID mismatch quirk for ASUS ROG Z13 Folio [+ + +]
Author: Lovekesh Solanki <[email protected]>
Date:   Wed Aug 5 01:50:31 2026 +0530

    HID: multitouch: Add report ID mismatch quirk for ASUS ROG Z13 Folio
    
    [ Upstream commit aaaea79efba5a27cb9e0a5a628046d829d3f2cbb ]
    
    Commit e716edafedad ("HID: multitouch: Check to ensure report
    responses match the request") introduced validating GET_FEATURE
    responses return the requested report ID.
    
    ASUS ROG Z13 Flow (2025) GZ302EA touchpad (USB 0b05:1a30) returns
    a different report ID for Win8 feature request. Before this check,
    the response was still processed and allowed device to switch into
    its full Touchpad Precision mode.
    
    After the validation, the response is discarded before
    hid_report_raw_event() processes it and device remains in fallback
    mode and no longer exposes ABS_MT_SLOT, ABS_MT_TOOL_TYPE or the
    multi-finger BTN_TOOL_* capabilities for palm rejection.
    
    Add a device quirk to allow the known firmware behavior for
    this device while preserving report ID validation for all other
    devices.
    
    The device previously matched the generic MT_CLS_WIN_8 entry, so
    base the new class on MT_CLS_WIN_8 to keep it on the same quirk set
    as before the regression.  MT_QUIRK_CONFIDENCE must be set
    explicitly: it is normally enabled by the class name check in
    mt_touch_input_mapping(), which only matches the MT_CLS_WIN_8*
    names, and it is what makes ABS_MT_TOOL_TYPE available for
    touchpads.
    
    Fixes: e716edafedad ("HID: multitouch: Check to ensure report responses match the request")
    
    Signed-off-by: Lovekesh Solanki <[email protected]>
    Reported-by: mayhemandcoffee <[email protected]>
    Closes: https://bugzilla.kernel.org/show_bug.cgi?id=221774
    Tested-by: mayhemandcoffee <[email protected]>
    Link: https://bugzilla.kernel.org/show_bug.cgi?id=221774
    Signed-off-by: Jiri Kosina <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

HID: quirks: add ALWAYS_POLL quirk for SDINNOVATION gaming keyboard [+ + +]
Author: Junjie Cao <[email protected]>
Date:   Mon Aug 24 11:14:19 2026 +0800

    HID: quirks: add ALWAYS_POLL quirk for SDINNOVATION gaming keyboard
    
    commit cdb669a3b8f844aca71fc3224990157d61562165 upstream.
    
    The SDINNOVATION gaming keyboard (USB ID 36ae:feab) stops reporting
    input events after its RGB lighting mode is switched about twice.
    Disabling USB autosuspend and unbinding the other HID interfaces make
    no difference; the issue does not occur on Windows.
    
    HID_QUIRK_ALWAYS_POLL alone resolves it, verified on 7.1.8 via
    usbhid.quirks=0x36ae:0xfeab:0x400.
    
    Reported-by: Marco Carvalho <[email protected]>
    Link: https://bugzilla.redhat.com/show_bug.cgi?id=2514627
    Cc: [email protected]
    Signed-off-by: Junjie Cao <[email protected]>
    Signed-off-by: Benjamin Tissoires <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

HID: wacom: fix OOB read in wacom_wac_pen_serial_enforce() [+ + +]
Author: Wei Jie LAW <[email protected]>
Date:   Wed Sep 9 09:15:13 2026 +0800

    HID: wacom: fix OOB read in wacom_wac_pen_serial_enforce()
    
    commit 9aa237cf66495b2426ddde8532e9b08a0ed83aaa upstream.
    
    The 'wacom_wac_pen_serial_enforce()' function may calculate and pass an
    invalid offset to hid_field_extract(), resulting in memory reads at
    incorrect addresses -- possibly beyond the end of the report.  If a
    field in the HID descriptor lists more usages than its Report Count
    actually reserves space for, the function's inner 'j' will walk past
    the end of the field:
    
            for (i = 0; i < report->maxfield; i++) {
                    for (j = 0; j < report->field[i]->maxusage; j++) {
                            ...
                            value = hid_field_extract(hdev, raw_data + 1,
                                                      offset + j * size, size);
    
    A descriptor listing 12288 usages against Report Count 1 has the loop
    extract the usage at index 12287 from bit offset 98296 -- about 12 KB
    past a 2-byte received report.  The value is stored in
    wacom_wac->serial[0] and can reach userspace as an MSC_SERIAL event,
    making this an information disclosure.
    
    Clamp the loop to field->report_count, the number of value slots the
    report holds.  Value slots past the last declared usage are still
    scanned; they reuse that usage (HID 1.11, 6.2.2.8).
    
    Verified on v6.12.105 with a UHID reproducer: a 2-byte report from
    such a descriptor trips KASAN before the patch and not after it.
    
    Fixes: 83417206427b ("HID: wacom: Queue events with missing type/serial data for later processing")
    Suggested-by: Jason Gerecke <[email protected]>
    Cc: [email protected]
    Assisted-by: Claude:claude-opus-5
    Assisted-by: GLM:glm-5.3
    Signed-off-by: Wei Jie Law <[email protected]>
    Reviewed-by: Jason Gerecke <[email protected]>
    Signed-off-by: Jiri Kosina <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

HID: winwing: fix use-after-free in force feedback teardown [+ + +]
Author: René Onier <[email protected]>
Date:   Wed Aug 12 13:13:59 2026 -0400

    HID: winwing: fix use-after-free in force feedback teardown
    
    [ Upstream commit a1a5ad37e50ceb192c07ac7e2d7143638cd4d110 ]
    
    winwing_init_ff() passes the driver's private data, allocated with
    devm_kzalloc() in winwing_probe(), as the effect context to
    input_ff_create_memless(). The memoryless force-feedback core takes
    ownership of that pointer and frees it with kfree() from
    input_ff_destroy() (ml_ff_destroy()) when the input device is
    destroyed.
    
    Freeing a devm-managed allocation with kfree() is an invalid free, and
    the same object is then released again by devres when the HID device is
    torn down, a double free. As the allocation also embeds the LED class
    devices, their timers and work item live on freed memory and the slab
    gets corrupted. This triggers on unbind, rmmod, hot-unplug and on system
    suspend, where the firmware cache walks the now-corrupt devres list.
    KASAN reports:
    
      BUG: KASAN: invalid-free in input_ff_destroy
      Allocated by task N:
        winwing_probe
    
    Pass NULL as the memless context instead and fetch the driver data from
    the input device in winwing_play_effect(): the HID core already stores
    the hid_device as the input device's drvdata. The force-feedback core
    then owns nothing that it must not free.
    
    Fixes: 42d020b54edc ("HID: winwing: Enable rumble effects")
    Signed-off-by: René Onier <[email protected]>
    Signed-off-by: Jiri Kosina <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
i2c: qcom-geni: Fix hardcoded clock index in SE_GENI_CLK_SEL [+ + +]
Author: Viken Dadhaniya <[email protected]>
Date:   Tue Sep 29 08:46:17 2026 -0400

    i2c: qcom-geni: Fix hardcoded clock index in SE_GENI_CLK_SEL
    
    [ Upstream commit cb97bf3d4f91453b881acaf8e9f0cc47bb40b604 ]
    
    qcom_geni_i2c_conf() writes a hardcoded 0 to SE_GENI_CLK_SEL, which
    selects an index from the hardware clock performance table. This always
    picks the first table entry regardless of the actual source clock
    configuration. On platforms where the matching entry is not at index 0,
    the wrong source clock divider is active and the I2C bus runs at an
    incorrect frequency.
    
    Use geni_se_clk_freq_match() in geni_i2c_clk_map_idx() to find the
    performance table index for the source clock (32 MHz or 19.2 MHz). Store
    the resolved index in a new clk_idx field in geni_i2c_dev and write it
    to SE_GENI_CLK_SEL instead of the hardcoded 0.
    
    Fixes: 37692de5d523 ("i2c: i2c-qcom-geni: Add bus driver for the Qualcomm GENI I2C controller")
    Signed-off-by: Viken Dadhaniya <[email protected]>
    Cc: <[email protected]> # v4.19+
    Reviewed-by: Mukesh Kumar Savaliya <[email protected]>
    Signed-off-by: Andi Shyti <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

i2c: qcom-geni: Isolate serial engine setup [+ + +]
Author: Praveen Talari <[email protected]>
Date:   Tue Sep 29 08:46:14 2026 -0400

    i2c: qcom-geni: Isolate serial engine setup
    
    [ Upstream commit d8d3bb127ad119853ddcf5da8f546cf37c3cc346 ]
    
    Moving the serial engine setup to geni_i2c_init() API for a cleaner
    probe function and utilizes the PM runtime API to control resources
    instead of direct clock-related APIs for better resource management.
    
    Enables reusability of the serial engine initialization like
    hibernation and deep sleep features where hardware context is lost.
    
    Signed-off-by: Praveen Talari <[email protected]>
    Acked-by: Viken Dadhaniya <[email protected]>
    Reviewed-by: Konrad Dybcio <[email protected]>
    Reviewed-by: Mukesh Kumar Savaliya <[email protected]>
    Tested-by: Mattijs Korpershoek <[email protected]>
    Signed-off-by: Andi Shyti <[email protected]>
    Link: https://lore.kernel.org/r/20260617-enable-i2c-on-sa8255p-v7-2-ad736dbeab57@oss.qualcomm.com
    Stable-dep-of: cb97bf3d4f91 ("i2c: qcom-geni: Fix hardcoded clock index in SE_GENI_CLK_SEL")
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

i2c: qcom-geni: Move resource initialization to separate function [+ + +]
Author: Praveen Talari <[email protected]>
Date:   Tue Sep 29 08:46:15 2026 -0400

    i2c: qcom-geni: Move resource initialization to separate function
    
    [ Upstream commit ed4b34033db25a0f35bb84289377e1916ffe2329 ]
    
    Refactor the resource initialization in geni_i2c_probe() by introducing
    a new geni_i2c_resources_init() function and utilizing the common
    geni_se_resources_init() framework and clock frequency mapping, making the
    probe function cleaner.
    
    Signed-off-by: Praveen Talari <[email protected]>
    Acked-by: Viken Dadhaniya <[email protected]>
    Reviewed-by: Konrad Dybcio <[email protected]>
    Tested-by: Mattijs Korpershoek <[email protected]>
    Signed-off-by: Andi Shyti <[email protected]>
    Link: https://lore.kernel.org/r/20260617-enable-i2c-on-sa8255p-v7-3-ad736dbeab57@oss.qualcomm.com
    Stable-dep-of: cb97bf3d4f91 ("i2c: qcom-geni: Fix hardcoded clock index in SE_GENI_CLK_SEL")
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

i2c: qcom-geni: release DMA channels on probe error [+ + +]
Author: Shengzhuo Wei <[email protected]>
Date:   Thu Aug 27 23:43:03 2026 +0800

    i2c: qcom-geni: release DMA channels on probe error
    
    commit 268aacb2e2a5c94d08e23194961234d0710c8407 upstream.
    
    geni_i2c_init() grabs exclusive GPI tx/rx DMA channels when the serial
    engine runs in GPI mode. If i2c_add_adapter() subsequently fails, probe
    returns without releasing the channels, because the remove callback is
    not invoked after a failed probe.
    
    The adapter-registration failure path used to release the channels via
    its err_dma label; that release was dropped when the probe tail was
    restructured into geni_i2c_init().
    
    Release the channels on the adapter-registration failure path, mirroring
    geni_i2c_remove().
    
    Fixes: d8d3bb127ad1 ("i2c: qcom-geni: Isolate serial engine setup")
    Assisted-by: GLM:5.3
    Signed-off-by: Shengzhuo Wei <[email protected]>
    Reviewed-by: Konrad Dybcio <[email protected]>
    Reviewed-by: Mukesh Kumar Savaliya <[email protected]>
    Signed-off-by: Andi Shyti <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

i2c: qcom-geni: Store of_device_id data in driver private struct [+ + +]
Author: Praveen Talari <[email protected]>
Date:   Tue Sep 29 08:46:16 2026 -0400

    i2c: qcom-geni: Store of_device_id data in driver private struct
    
    [ Upstream commit 692e0c84db5fdd88c242eadc873d498787c94e3e ]
    
    To avoid repeatedly fetching and checking platform data across various
    functions, store the struct of_device_id data directly in the i2c
    private structure. This change enhances code maintainability and reduces
    redundancy.
    
    Signed-off-by: Praveen Talari <[email protected]>
    Acked-by: Viken Dadhaniya <[email protected]>
    Reviewed-by: Konrad Dybcio <[email protected]>
    Tested-by: Mattijs Korpershoek <[email protected]>
    Signed-off-by: Andi Shyti <[email protected]>
    Link: https://lore.kernel.org/r/20260617-enable-i2c-on-sa8255p-v7-5-ad736dbeab57@oss.qualcomm.com
    Stable-dep-of: cb97bf3d4f91 ("i2c: qcom-geni: Fix hardcoded clock index in SE_GENI_CLK_SEL")
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
ip6_gre: Call ip6erspan_tunnel_unlink_md() in ip6erspan_changelink(). [+ + +]
Author: Kuniyuki Iwashima <[email protected]>
Date:   Wed Sep 16 23:09:24 2026 +0000

    ip6_gre: Call ip6erspan_tunnel_unlink_md() in ip6erspan_changelink().
    
    [ Upstream commit dd47bcf279f1083f09bf5266890b26263361022b ]
    
    The cited commit accidentally added ip6gre_tunnel_unlink_md()
    in ip6erspan_changelink().
    
    Let's correct it to ip6erspan_tunnel_unlink_md().
    
    Fixes: b80d0b93b991 ("net: ip6_gre: fix tunnel metadata device sharing.")
    Signed-off-by: Kuniyuki Iwashima <[email protected]>
    Reviewed-by: Xuanqiang Luo <[email protected]>
    Reviewed-by: Ido Schimmel <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
ip_gre: Reject enabling collect metadata through changelink [+ + +]
Author: Xuanqiang Luo <[email protected]>
Date:   Mon Sep 21 11:18:59 2026 +0800

    ip_gre: Reject enabling collect metadata through changelink
    
    [ Upstream commit a3f315be9d30eeb6938d11fa17fd4b32d52f7c42 ]
    
    ipgre_netlink_parms() can enable collect_md on an existing GRE, GRETAP
    or ERSPAN device. Unlike newlink, changelink does not enforce metadata
    tunnel uniqueness. Converting a non-metadata device can therefore
    replace the metadata receive entry for another device of the same type
    in the same netns. Deleting either device then clears the shared entry,
    breaking metadata receive lookup for the surviving device.
    
    If parameter validation fails after collect_md is set, deleting the
    modified device can also clear an entry it never owned.
    
    Reject enabling metadata mode in both changelink callbacks before any
    encapsulation or tunnel parameters are modified. Allow requests that
    repeat the metadata attribute on an existing metadata device.
    
    Fixes: 2e15ea390e6f ("ip_gre: Add support to collect tunnel metadata.")
    Signed-off-by: Xuanqiang Luo <[email protected]>
    Reviewed-by: Ido Schimmel <[email protected]>
    Reviewed-by: Hangbin Liu <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
ipe: fix use-after-free when auditing a newly loaded policy [+ + +]
Author: Fan Wu <[email protected]>
Date:   Tue Sep 22 20:13:48 2026 -0700

    ipe: fix use-after-free when auditing a newly loaded policy
    
    commit 9814077275eca36ebf8d510d2f076d235ff9f51a upstream.
    
    new_policy() audits the policy after ipe_new_policyfs_node() publishes it
    and drops the new directory's inode lock. A concurrent delete can free
    the policy while ipe_audit_policy_load() is still using it.
    
    Audit the successful load under that lock.
    
    Fixes: f44554b5067b ("audit,ipe: add IPE auditing support")
    Cc: [email protected]
    Assisted-by: LLM
    [FW: remove model name according to latest guideline]
    Signed-off-by: Fan Wu <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

ipe: protect the dm-verity root hash with RCU [+ + +]
Author: Fan Wu <[email protected]>
Date:   Tue Sep 22 20:13:49 2026 -0700

    ipe: protect the dm-verity root hash with RCU
    
    commit 2776e9c28513a1c855a94792b292cbcc533418c8 upstream.
    
    ipe_bdev_setintegrity() frees the old root hash when dm-verity publishes
    a new one on ->preresume, while policy evaluation can still be
    dereferencing it.
    
    Protect the root hash with RCU. The evaluation path already runs under
    rcu_read_lock().
    
    Fixes: e155858dd995 ("ipe: add support for dm-verity as a trust provider")
    Cc: [email protected]
    Assisted-by: LLM
    [FW: remove model name according to latest guideline]
    Signed-off-by: Fan Wu <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
ipv4: fib: fix data-race and stale genid check around nh->nh_saddr [+ + +]
Author: Linkui Xiao <[email protected]>
Date:   Wed Sep 16 20:53:16 2026 +0800

    ipv4: fib: fix data-race and stale genid check around nh->nh_saddr
    
    [ Upstream commit 46bc52d13594848023e681860df8700c8db14354 ]
    
    fib_select_multipath() compares nexthop_nh->nh_saddr against the flow
    source address with no lock held, while fib_info_update_nhc_saddr()
    stores a new value from another CPU as soon as the preferred source
    address of the egress device changes.
    
    Commit 195374d89368 ("ipv4: fib: annotate races around nh->nh_saddr_genid
    and nh->nh_saddr") added WRITE_ONCE() on the store side and READ_ONCE()
    in fib_result_prefsrc() after syzbot reported
    
            BUG: KCSAN: data-race in fib_select_path / fib_select_path
    
    but it only covered that reader. fib_select_multipath(), reached from
    fib_select_path(), is a second lockless reader of nh->nh_saddr and was
    left bare.
    
    Moreover, nh_saddr is only meaningful when nh_saddr_genid matches
    dev_addr_genid, as established by commit 436c3b66ec98 ("ipv4: Invalidate
    nexthop cache nh_saddr more correctly."). fib_select_multipath()
    skips that validation, so it can score a nexthop using a stale source
    address and skew the ECMP selection.
    
    Annotate both reads with READ_ONCE() and refresh the cached source
    address via fib_info_update_nhc_saddr() when the genid does not match,
    mirroring fib_result_prefsrc().
    
    Fixes: 32607a332cfe ("ipv4: prefer multipath nexthop that matches source address")
    Signed-off-by: Linkui Xiao <[email protected]>
    Reviewed-by: Ido Schimmel <[email protected]>
    Reviewed-by: Eric Dumazet <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
ipv6: do not let ipv6_find_hdr() return an offset past the packet end [+ + +]
Author: Norbert Szetei <[email protected]>
Date:   Wed Sep 16 21:57:53 2026 +0200

    ipv6: do not let ipv6_find_hdr() return an offset past the packet end
    
    commit ee319bd3a0e976af5087cbe59ebc50a66f31d202 upstream.
    
    ipv6_find_hdr() walks the extension header chain, skipping each header by
    the length that header itself declares.  ipv6_optlen() returns up to 2048,
    and the skip is never checked against skb->len, so the offset stored in
    *offset can point past the end of the packet.
    
    openvswitch installs that offset as the transport header, and
    update_ipv6_checksum() then reads and writes the transport checksum field
    out of bounds:
    
      BUG: KASAN: slab-use-after-free in inet_proto_csum_replace16+0x445/0x470
      Read of size 2 at addr ffff88810b754b06 by task ovs_ipv6_oob/629
      CPU: 4 UID: 1000 PID: 629 Comm: ovs_ipv6_oob Tainted: G N 7.3.0-rc3+ #348
      Call Trace:
       inet_proto_csum_replace16+0x445/0x470
       set_ipv6_addr+0x3dd/0x460
       do_execute_actions+0x6a3d/0x7c40
       ovs_execute_actions+0xfd/0x480
       ovs_packet_cmd_execute+0xc38/0xf20
       genl_rcv_msg+0x59e/0x870
       netlink_rcv_skb+0x18b/0x450
       genl_rcv+0x2d/0x40
       netlink_unicast+0x6bc/0xa20
    
      The buggy address belongs to the object at ffff88810b754980
       which belongs to the cache skbuff_small_head of size 704
      The buggy address is located 390 bytes inside of
       freed 704-byte region [ffff88810b754980, ffff88810b754c40)
    
    Other callers use that offset too, so bound it here rather than in one
    caller.
    
    Reject a header whose declared length does not fit in the packet.
    ipv6_find_hdr() already fails with -EBADMSG on a malformed chain, so this
    adds no new failure mode.
    
    Fixes: f8f626754ebe ("ipv6: Move ipv6_find_hdr() out of Netfilter code.")
    Suggested-by: Ilya Maximets <[email protected]>
    Suggested-by: Eric Dumazet <[email protected]>
    Cc: [email protected]
    Signed-off-by: Norbert Szetei <[email protected]>
    Reviewed-by: Ido Schimmel <[email protected]>
    Reviewed-by: Ilya Maximets <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

ipv6: Fix dst leak for uncached routes. [+ + +]
Author: Kuniyuki Iwashima <[email protected]>
Date:   Sun Sep 20 19:14:32 2026 +0000

    ipv6: Fix dst leak for uncached routes.
    
    [ Upstream commit be31fe6333f534155e6b408f1ef6d77974bb41aa ]
    
    ip6_route_output_flags(), ip6_rt_put_flags(), and ip6_dst_check()
    detect an uncached route by list_empty(&rt->dst.rt_uncached),
    which replaced the static DST_NOCACHE flag check in commit
    a4c2fd7f7891 ("net: remove DST_NOCACHE flag").
    
    When a device is unregistered, rt6_uncached_list_flush_dev()
    unlinks uncached routes tied to the device from rt6_uncached_list.
    
    Previously, they were moved to another list with list_move()
    (__list_del_entry() + list_add()), and since commit 98aa546af5e4
    ("inet: remove (struct uncached_list)->quarantine"), the routes
    are just unlinked with list_del_init().
    
    If list_del_init() runs concurrently, list_empty() evaluates to
    true; ip6_route_output_flags() calls dst_hold_safe() incorrectly
    and ip6_rt_put_flags() skips ip6_rt_put(), leaking dst, and thus
    dev tied via rt->from as well.
    
    The same race is partially fixed by commit 9a6f0c4d5796 ("dst:
    fix races in rt6_uncached_list_del() and rt_del_uncached_list()").
    
    Let's check rt6->dst.rt_uncached_list instead.
    
    Note that IPv4 does not have the same issue.
    
    Fixes: 98aa546af5e4 ("inet: remove (struct uncached_list)->quarantine")
    Signed-off-by: Kuniyuki Iwashima <[email protected]>
    Reviewed-by: Hangbin Liu <[email protected]>
    Reviewed-by: Xuanqiang Luo <[email protected]>
    Reviewed-by: Ido Schimmel <[email protected]>
    Reviewed-by: Eric Dumazet <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

ipv6: Prevent rt6_insert_exception() for dying fib6_info. [+ + +]
Author: Kuniyuki Iwashima <[email protected]>
Date:   Fri Sep 18 08:22:05 2026 +0000

    ipv6: Prevent rt6_insert_exception() for dying fib6_info.
    
    [ Upstream commit 0346ec2f080b40d95ed05b853bb9226289e75212 ]
    
    Before the cited commit, fib6_nh_flush_exceptions() always set
    from->exception_bucket_flushed = 1 under rt6_exception_lock to
    prevent rt6_insert_exception() from inserting a new exception
    for a dying fib6_info.
    
    The flag was replaced with the FIB6_EXCEPTION_BUCKET_FLUSHED
    bit stored in nh->rt6i_exception_bucket.
    
    The problem is that now the bit is only set when the bucket
    is not NULL and fib6_nh_flush_exceptions() is called from
    fib6_nh_release() after fib6_ref has already reached zero.
    
    If rt6_insert_exception() is called while the target fib6_info
    is being removed via fib6_purge_rt(), a new exception could be
    created successfully because rt6_flush_exceptions() no longer
    sets the bit.
    
    This creates a reference cycle between the fib6_info and the
    exception route, leaking the fib6_info, its nexthop device,
    and all per-CPU routes in fib6_nh->rt6i_pcpu, which stalls netdev
    unregistration.
    
    [   34.680602] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
    [   44.920675] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
    [   55.176582] unregister_netdevice: waiting for gre6 to become free. Usage count = 68
    
    Let's call fib6_drop_pcpu_from() before rt6_flush_exceptions(),
    to set fib6_destroying before rt6_exception_lock, and check
    f6i->fib6_destroying in rt6_insert_exception().
    
    Note that FIB6_EXCEPTION_BUCKET_FLUSHED logic is dead and
    we can clean it up in net-next.
    
    Fixes: cc5c073a693f ("ipv6: Move exception bucket to fib6_nh")
    Signed-off-by: Kuniyuki Iwashima <[email protected]>
    Reviewed-by: Ido Schimmel <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

ipv6: sr: enforce exact attribute length for SEG6_ATTR_DST [+ + +]
Author: Hui Peng <[email protected]>
Date:   Mon Sep 21 04:40:25 2026 +0000

    ipv6: sr: enforce exact attribute length for SEG6_ATTR_DST
    
    commit 2d959c75c27f90e9ec489d18ce5ee6b852ad4741 upstream.
    
    In seg6_genl_policy, SEG6_ATTR_DST is defined with .type = NLA_BINARY and
    .len = sizeof(struct in6_addr). For NLA_BINARY, .len only enforces the
    maximum payload length and permits shorter payloads (e.g., 0 bytes).
    When seg6_genl_set_tunsrc() copies sizeof(struct in6_addr) bytes via
    kmemdup(val, sizeof(*val), GFP_KERNEL), a short SEG6_ATTR_DST attribute
    triggers a 16-byte out-of-bounds read past skb->tail into uninitialized
    skb->head memory, which is stored in sdata->tun_src and leaked back to
    userspace via SEG6_CMD_GET_TUNSRC.
    
    Switch SEG6_ATTR_DST in seg6_genl_policy to
    NLA_POLICY_EXACT_LEN(sizeof(struct in6_addr)) so that generic netlink
    validation rejects any attribute whose length is not exactly
    sizeof(struct in6_addr) with -ERANGE.
    
    Tested in QEMU against Linux 7.3.0-rc3 by sending a SEG6_CMD_SET_TUNSRC
    Generic Netlink message with a 0-byte SEG6_ATTR_DST attribute followed
    by SEG6_CMD_GET_TUNSRC. On the unfixed kernel, SEG6_CMD_SET_TUNSRC
    succeeds (err = 0) and SEG6_CMD_GET_TUNSRC leaks 16 bytes of
    uninitialized kernel heap memory (tun_src =
    836a61ecc4d25a1042a8d60411cfb378); with this patch applied,
    SEG6_CMD_SET_TUNSRC is rejected by netlink policy validation with
    -ERANGE (-34) and tun_src remains zeroed.
    
    Fixes: 915d7e5e5930 ("ipv6: sr: add code base for control plane support of SR-IPv6")
    Cc: [email protected]
    Signed-off-by: Hui Peng <[email protected]>
    Reviewed-by: Hangbin Liu <[email protected]>
    Reviewed-by: Justin Iurman <[email protected]>
    Reviewed-by: Andrea Mayer <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
ipvs: revalidate ihl before icmp_send [+ + +]
Author: Julian Anastasov <[email protected]>
Date:   Fri Sep 11 14:43:15 2026 +0300

    ipvs: revalidate ihl before icmp_send
    
    [ Upstream commit e290145564886d6a3038810c621f738c1fe9fa51 ]
    
    While the outer IP header is already pulled into the skb head, we must
    be careful and revalidate the embedded headers after reading them from
    the skb frags to prevent possible out-of-bounds access.
    
    One such place reported by Sashiko is ip_vs_in_icmp() where local
    process can change the ihl field and after pskb_may_pull() we can see
    larger value. Even if icmp_send() has checks to prevent out-of-bounds
    access, play safe and add check to drop the packet if the ihl field is
    changed.  As the outer headers are pulled, make sure the transport
    header is updated too, it was used before commit 7fcc2fe39fed ("net:
    icmp: avoid invalid transport header access in icmp_send tracepoint")
    
    Fixes: f2edb9f7706d ("ipvs: implement passive PMTUD for IPIP packets")
    Link: https://sashiko.dev/#/patchset/20260806105211.34622-1-ja%40ssi.bg
    Signed-off-by: Julian Anastasov <[email protected]>
    Signed-off-by: Pablo Neira Ayuso <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
kprobes: Fix permanent hang when flushing the kprobe optimizer [+ + +]
Author: Andrea Parri <[email protected]>
Date:   Thu Sep 24 11:21:39 2026 +0200

    kprobes: Fix permanent hang when flushing the kprobe optimizer
    
    commit 5bfa9f1a9dcb6ecb607adbc1c0226605c972935b upstream.
    
    Writing 0 to /proc/sys/debug/kprobes-optimization while a kprobe is
    jump-optimized never returns. The writer sleeps in D state forever with
    kprobe_sysctl_mutex held, so any later read or write of that sysctl
    hangs as well. For example, with vfs_read+9 as an optimizable address
    in this build:
    
      # cd /sys/kernel/tracing
      # echo 'p:myprobe vfs_read+9' >> kprobe_events
      # echo 1 > events/kprobes/myprobe/enable
      # # wait until /sys/kernel/debug/kprobes/list shows [OPTIMIZED]
      # echo 0 > /proc/sys/debug/kprobes-optimization
    
      INFO: task sh:246 blocked for more than 10 seconds.
      Call Trace:
       <TASK>
       __schedule+0x1176/0x4f70
       schedule+0xdc/0x2c0
       schedule_timeout+0x17b/0x260
       wait_for_completion+0x173/0x3c0
       wait_for_kprobe_optimizer_locked+0xbc/0x130
       proc_kprobes_optimization_handler+0x156/0x1b0
       proc_sys_call_handler+0x324/0x490
       vfs_write+0x52d/0xfe0
       ksys_write+0xff/0x200
       do_syscall_64+0x106/0x630
       entry_SYSCALL_64_after_hwframe+0x77/0x7f
       </TASK>
      ...
      INFO: task cat:265 is blocked on a mutex likely owned by task sh:246.
    
    wait_for_kprobe_optimizer_locked() reinitializes optimizer_completion,
    asks the optimizer thread to flush and sleeps in wait_for_completion().
    The thread drains the (un)optimizing lists, but calls complete() only
    if completion_done() is true, i.e. if the completion is already done,
    which never happens while someone waits. disarm_all_kprobes() and
    kprobe_trace_self_tests_init() wait the same way.
    
    Calling complete() unconditionally would not be enough: the waiter
    drops kprobe_mutex while it sleeps, and nothing else serializes the
    sysctl handler against the debugfs "enabled" file. A second flusher
    that still finds the lists non-empty, e.g. because a disabled probe is
    queued for unoptimizing, reinitializes the completion under the first:
    
      sysctl write                      debugfs "enabled" write
      unoptimize_all_kprobes()
        wait_for_kprobe_optimizer_locked()
          init_completion(c)
          mutex_unlock(&kprobe_mutex)
          wait_for_completion(c)
                                        disarm_all_kprobes()
                                          wait_for_kprobe_optimizer_locked()
                                            init_completion(c)
                                              // c->wait is reset, the first
                                              // waiter is off the queue
                                            mutex_unlock(&kprobe_mutex)
                                            wait_for_completion(c)
      kprobe_optimizer()
        complete(c)
          // wakes the debugfs writer only
    
    where c is &optimizer_completion. Lining up the two writes during an
    optimizer pass loses the sysctl writer this way.
    
    Replace the completion with a counter of optimizer passes, bumped at the
    end of each pass and signalled with wake_up_var_locked(), both under
    kprobe_mutex. A flusher samples the count and waits with
    wait_var_event_mutex(), which drops kprobe_mutex only while sleeping, so
    a new count means a whole pass ran in the meantime. Nothing is
    reinitialized, so several flushers can sleep in the wait at once.
    
    Link: https://lore.kernel.org/all/[email protected]/
    
    Fixes: 73c12f209462 ("kprobes: Use dedicated kthread for kprobe optimizer")
    Cc: [email protected]
    Assisted-by: LLM
    Signed-off-by: Andrea Parri <[email protected]>
    Signed-off-by: Masami Hiramatsu (Google) <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2 [+ + +]
Author: Fuad Tabba <[email protected]>
Date:   Tue Aug 25 09:59:48 2026 +0100

    KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2
    
    [ Upstream commit 4f16c5fc8dc4c5596e3777ab9f449a54e3f85fd5 ]
    
    pkvm_init_features_from_host() takes KVM_ARCH_FLAG_GUEST_HAS_SVE and
    KVM_ARM_VCPU_SVE from the host separately, but pkvm_vcpu_init_sve()
    tests the bit while vcpu_has_sve() reads the flag. A host that sets the
    flag without the bit gets a vCPU with a NULL sve_state that the world
    switch loads the guest's SVE state from.
    
    Derive the flag from the bit, and drop the protected path's copy of the
    host's flag, which is dead code since protected VMs are not allowed SVE.
    
    Fixes: 41d6028e28bd ("KVM: arm64: Convert the SVE guest vcpu flag to a vm flag")
    Signed-off-by: Fuad Tabba <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: arm64: Do not clear VM-wide SVE feature on vCPU init failure [+ + +]
Author: Fuad Tabba <[email protected]>
Date:   Tue Aug 25 09:59:46 2026 +0100

    KVM: arm64: Do not clear VM-wide SVE feature on vCPU init failure
    
    [ Upstream commit a1b3c788ad31837e348075e93dbba3f447492776 ]
    
    pkvm_vcpu_init_sve() clears KVM_ARM_VCPU_SVE in kvm->arch.vcpu_features
    when it fails, but vcpu_has_sve() tests KVM_ARCH_FLAG_GUEST_HAS_SVE,
    which is left set. Later vCPUs on that VM then skip the SVE setup and
    register with a NULL sve_state, which the guest's first FP access hands
    to sve_load_state().
    
    Return the error without touching vcpu_features.
    
    Fixes: 5db1bef93342 ("KVM: arm64: Track SVE state in the hypervisor vcpu structure")
    Signed-off-by: Fuad Tabba <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode [+ + +]
Author: Fuad Tabba <[email protected]>
Date:   Mon Sep 14 10:38:38 2026 +0100

    KVM: arm64: Don't WARN on an unknown VM ioctl in protected mode
    
    commit 49d9d295d69d07e850cae35933ba8519e2915f26 upstream.
    
    kvm_pkvm_ioctl_allowed() WARNs when kvm_get_cap_for_kvm_ioctl() doesn't
    find the ioctl number in vm_ioctl_caps[], and kvm_arch_vm_ioctl() calls
    it for every number the generic code doesn't handle, so
    ioctl(vm_fd, 0xdeadbeef) from userspace taints a pKVM host and panics it
    under panic_on_warn. The lookup is fed userspace input: return false,
    and userspace gets the -EINVAL kvm_arch_vm_ioctl() returns for that
    number on a host without pKVM.
    
    Fixes: b12b3b04f6ba0 ("KVM: arm64: Check whether a VM IOCTL is allowed in pKVM")
    Cc: [email protected]
    Signed-off-by: Fuad Tabba <[email protected]>
    Reviewed-by: Suzuki K Poulose <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: arm64: Fix AArch32 DBGBXVR handling [+ + +]
Author: Karl Mehltretter <[email protected]>
Date:   Mon Aug 10 02:56:16 2026 +0200

    KVM: arm64: Fix AArch32 DBGBXVR<n> handling
    
    commit 6b1bca1b1ab77f60a62087337bfe6e2f0efb9e6d upstream.
    
    The consolidation of the breakpoint and watchpoint register accessors
    switched DBGBXVR<n> from trap_bvr() to trap_dbg_wb_reg(). The latter
    selects backing storage with demux_wb_reg(), which only handles Op2 values
    4 through 7. Since DBGBXVR<n> uses Op2 1, an AArch32 guest access hits
    KVM_BUG_ON() and marks the VM dead.
    
    DBGBXVR<n> aliases DBGBVR<n>_EL1[63:32], and its AA32(HI) descriptor
    already selects the upper half. Map Op2 1 to dbg_bvr[] alongside Op2 4,
    restoring the pre-regression behavior.
    
    Fixes: 3ce9f3357e9e ("KVM: arm64: Fold DBGxVR/DBGxCR accessors into common set")
    Cc: [email protected]
    Assisted-by: Claude:claude-opus-5
    Signed-off-by: Karl Mehltretter <[email protected]>
    Reviewed-by: Marc Zyngier <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP [+ + +]
Author: Mark Brown <[email protected]>
Date:   Tue Sep 1 22:47:00 2026 +0100

    KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP
    
    [ Upstream commit 089e4f3c4862ba3f29dff2361caa8084879194fd ]
    
    The encoding to trap mapping currently maps a FGT on OP_GCSPOPX to
    HFGITR_EL2.nGCSEPP but as per DDI0601 2026-06 this FGT controls trapping
    of GCSPUSHX and GCSPOPCX, and not the separate GCSPOPX instruction.
    Update the mapping to reflect the architecture.
    
    Fixes: 863ac38984a82 ("KVM: arm64: Add missing HFGITR_EL2 FGT entries to nested virt")
    Reviewed-by: Leonardo Bras <[email protected]>
    Signed-off-by: Mark Brown <[email protected]>
    Reviewed-by: Lorenzo Stoakes (ARM) <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: arm64: Fix spurious warning for benign stage 2 teardown race [+ + +]
Author: Lorenzo Stoakes (ARM) <[email protected]>
Date:   Tue Sep 1 18:28:59 2026 +0100

    KVM: arm64: Fix spurious warning for benign stage 2 teardown race
    
    commit 38b70fc453c3112f1a62583b89903ae41116cc27 upstream.
    
    kvmtool was used to establish an L1 guest with 8 CPUs and 8 GiB of RAM, an
    L2 guest with 4 CPUs and 4 GiB of RAM and an L3 guest with 2 CPUs and 2 GiB
    of RAM, all of which was then exited.
    
    Under memory pressure in the L0 host warnings were observed due to
    migration triggered by compaction:
    
    WARNING: arch/arm64/kvm/mmu.c:336 at __unmap_stage2_range+0x64/0x80,
    CPU#5: kcompactd0/66
    
    Which was, in turn, triggered by an MMU notifier for the host invalidation:
    
    mmu_notifier_invalidate_range_start()
      -> ... -> kvm_mmu_notifier_invalidate_range_start()
        -> kvm_mmu_unmap_gfn_range()
          -> kvm_unmap_gfn_range()
            -> kvm_nested_s2_unmap()
              -> kvm_stage2_unmap_range()
                -> __unmap_stage2_range()
                   -> stage2_apply_range()
                   <- -EINVAL, triggering a WARN_ON()
    
    Racing with L0's teardown of stage 2 page tables:
    
    exit_mm()
      -> mmput()
        -> __mmput()
          -> exit_mmap()
            -> mmu_notifier_release()
              -> ... -> kvm_mmu_notifier_release()
                -> kvm_flush_shadow_all()
                  -> kvm_arch_flush_shadow_all()
                    -> kvm_free_stage2_pgd()
                      -> [ acquire kvm->mmu_lock for write ]
                      -> mmu->pgt = NULL [ among other tasks ]
                      -> [ release kvm->mmu_lock for write ]
    
    It turns out there is a benign race resulting in a spurious warning:
    
            Thread A - notify: migration   | Thread B - notify: release
            -------------------------------|---------------------------------
            < kvm->mmu_lock held >         |
            stage2_apply_range()           |
              get mmu->pgt, check !NULL    |
              ...                          | kvm_arch_flush_shadow_all()
              cond_resched_rwlock_write(); |   < contend, sleep kvm->mmu_lock >
            < drop kvm->mmu_lock >         |   < acquire kvm->mmu_lock>
                                           |   ...
                                           |   kvm_free_stage2_pgd()
                                           |     mmu->pgt = NULL
                                           |   < invalidate MMU >
                                           |   ...
                                           |   < release kvm->mmu_lock >
            [ scheduled ]                  |
            stage2_apply_range()           |
              < loop to next >             |
              get, mmu->pgt, check !NULL   |
              is NULL, return -EINVAL      |
            __unmap_stage2_range()         |
              WARN_ON(-EINVAL) <--- entirely spurious - the race was handled
                                     correctly.
    
    Fix the spurious warning by updating stage2_apply_range() to no longer
    treat concurrent PGT teardown on lock release as an error - whether the
    walker is tearing down page tables or doing something else this is a
    legitimate reason to abort the operation without error.
    
    This keeps the warning in place for all other circumstances.
    
    In practice only __unmap_stage2_range() actually does anything with the
    error so this only impacts that.
    
    Fixes: ec14c272408a ("KVM: arm64: nv: Unmap/flush shadow stage 2 page tables")
    Cc: [email protected]
    Reviewed-by: Yuan Yao <[email protected]>
    Reviewed-by: Marc Zyngier <[email protected]>
    Signed-off-by: Lorenzo Stoakes (ARM) <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: arm64: Match hyp text by physical address in fix_host_ownership() [+ + +]
Author: Fuad Tabba <[email protected]>
Date:   Tue Sep 8 12:07:11 2026 +0100

    KVM: arm64: Match hyp text by physical address in fix_host_ownership()
    
    [ Upstream commit 5a8b505ede133fb30ca3b3a19d0db00c08237615 ]
    
    On a non-hVHE host, fix_host_ownership_walker()'s test for PAGE_HYP_EXEC
    never matches: KVM_PGTABLE_PROT_UX is cleared at map time and only PX is
    reported on read-back. Hyp text is therefore donated rather than left
    read-only in the host stage-2, and the instruction dump in
    nvhe_hyp_panic_handler() reads a page the host has no access to.
    
    Match the text by physical address instead, in a helper a later patch
    reuses. A test on the permissions would leave any other executable
    mapping host-readable too.
    
    Fixes: 80cbfd7174f31 ("KVM: arm64: Honor UX/PX attributes for EL2 S1 mappings")
    Signed-off-by: Fuad Tabba <[email protected]>
    Reviewed-by: Vincent Donnefort <[email protected]>
    Tested-by: Vincent Donnefort <[email protected]>
    Reviewed-by: Marc Zyngier <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction [+ + +]
Author: Marc Zyngier <[email protected]>
Date:   Wed Sep 30 08:45:32 2026 -0400

    KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction
    
    [ Upstream commit e5843f4effaa2ffac3e789ecd4456403564961d4 ]
    
    We free the shadow S2 structures from kvm_arch_flush_shadow_all(), which
    is a Bad Idea(tm). Freeing the page tables is fair game (this is what
    this callback is for), but freeing the container that could still be
    referenced by another part of the system is not great.
    
    Instead, grow separate destructors that gets called when we tear the VM
    down for good. From there, we can nuke both the individual MMUs as well
    as the global array that points to them, safe in the knowledge that the
    vcpus themselves have been destroyed already.
    
    Fixes: 4f128f8e1aaac ("KVM: arm64: nv: Support multiple nested Stage-2 mmu structures")
    Reviewed-by: Lorenzo Stoakes (ARM) <[email protected]>
    Signed-off-by: Marc Zyngier <[email protected]>
    Cc: [email protected]
    Reviewed-by: Wei-Lin Chang <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: arm64: nv: Fix life cycle of the nested_mmus array [+ + +]
Author: Marc Zyngier <[email protected]>
Date:   Wed Sep 30 08:45:31 2026 -0400

    KVM: arm64: nv: Fix life cycle of the nested_mmus array
    
    [ Upstream commit 33346f8960c7bb6a3b4e273b5cfe25c5a8be349f ]
    
    The nested_mmus array holds the shadow page tables that are used when
    a guest is running a nested context. These structures are allocated on
    VCPU_INIT for whole guest, which implies that they may have to be
    relocated as the array grows.
    
    Should a VCPU_INIT occur whilst a vcpu is actively running an L2 and
    that the allocation requires relocation, that vcpu will still be
    running with a pointer to the previous structure, which will have been
    freed.
    
    Fix this by turning the array of structures to an array of pointers,
    which is now allocated at VM creation, sized to the absolute maximum
    that KVM can handle.
    
    In turn, each VCPU_INIT contributes S2_MMU_PER_VCPU to the pool. No
    reallocation is ever performed, and the life cycle of each object is
    much clearer:
    
    - the nested_mmus array is allocated in kvm_init_nested(), and freed
      in kvm_arch_destroy_vm()
    
    - s2_mmu structures are allocated in kvm_vcpu_init_nested(), and freed
      on kvm_arch_flush_shadow_all()
    
    Finally, the freeing of vcpu->arch.vncr_array is made consistent
    rather than being done on some failure paths, but not others.
    
    Fixes: 4f128f8e1aaa ("KVM: arm64: nv: Support multiple nested Stage-2 mmu structures")
    Reported-by: Shen Yongchao <[email protected]>
    Reported-by: Karl Mehltretter <[email protected]>
    Suggested-by: Karl Mehltretter <[email protected]>
    Acked-by: Lorenzo Stoakes (ARM) <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Marc Zyngier <[email protected]>
    Cc: [email protected]
    Reviewed-by: Wei-Lin Chang <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    
    [ Backport to 7.2: omit vncr_tlb_count initialization because this tree
      removed the counter in 87c2bbf189829 and does not have the subsequent
      VNCR TLB tracking reintroduction. Retain kvcalloc() for the per-vCPU
      MMU block, with the new fixed S2_MMU_PER_VCPU allocation size, because
      this call site has not undergone the upstream allocator conversion.
      Preserve the pointer-array lifetime changes required by e5843f4effaa2
      ("KVM: arm64: nv: Delay freeing of shadow S2 structures until VM
      destruction"). No functions are added. ]
    
    Stable-dep-of: e5843f4effaa ("KVM: arm64: nv: Delay freeing of shadow S2 structures until VM destruction")
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, teardown race [+ + +]
Author: Lorenzo Stoakes (ARM) <[email protected]>
Date:   Tue Sep 1 18:29:00 2026 +0100

    KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, teardown race
    
    commit 4c74e233cdedd11592775fae2a6243e67ca3f891 upstream.
    
    Commit 7270cc9157f4 ("KVM: arm64: nv: Handle VNCR_EL2 invalidation from MMU
    notifiers") introduced VNCR_EL2 invalidation in both kvm_nested_s2_unmap()
    and kvm_nested_s2_wp().
    
    However at the point of this being performed concurrent stage 2 teardown of
    a nested guest can cause kvm->arch.mmu.pgt to be set to NULL.
    
    This happens in kvm_flush_shadow_all() -> kvm_arch_flush_shadow_all() ->
    kvm_free_stage2_pgd() and is performed under the kvm->mmu_lock.
    
    Commit ec14c272408a ("KVM: arm64: nv: Unmap/flush shadow stage 2 page
    tables") introduced the teardown of the entire nested MMU range, which then
    invokes stage2_apply_range() with resched=true:
    
    mmu_notifier_invalidate_range_start()
      -> ... -> kvm_mmu_notifier_invalidate_range_start()
        -> kvm_mmu_unmap_gfn_range()
          -> kvm_unmap_gfn_range()
            -> kvm_nested_s2_unmap()
              -> kvm_stage2_unmap_range()
                -> __unmap_stage2_range()
                    -> stage2_apply_range()
    
    This means that stage2_apply_range() can drop the kvm->mmu_lock and thus
    concurrent progress can be made in lockstep with
    kvm_arch_flush_shadow_all().
    
    If kvm_arch_flush_shadow_all() advances ahead of stage2_apply_range() and
    completes its operation it guarantees a NULL pointer deref.
    
    Since kvm_free_stage2_pgd() is performed under the kvm->mmu_lock this will
    either be observed NULL or not and serialised against
    kvm_free_stage2_pgd().
    
    Resolve the issue by abstracting the invalidation to a new function,
    kvm_invalidate_vncr_ipa_all(), and check that the pgt is non-NULL before
    dereferencing it.
    
    Fixes: 7270cc9157f4 ("KVM: arm64: nv: Handle VNCR_EL2 invalidation from MMU notifiers")
    Cc: [email protected]
    Reviewed-by: Marc Zyngier <[email protected]>
    Signed-off-by: Lorenzo Stoakes (ARM) <[email protected]>
    Tested-by: Jonathan Davies <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0 [+ + +]
Author: Karl Mehltretter <[email protected]>
Date:   Sat Aug 29 07:48:55 2026 +0200

    KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0
    
    [ Upstream commit 64dc6f1db7e620f2e9337bb181f305fb0561da79 ]
    
    kvm_smccc_set_filter() only rejects a range if its inclusive end,
    base + nr_functions - 1, is below base. That catches an empty range
    (nr_functions == 0) at every nonzero base, but at base 0 the end wraps
    to U32_MAX and KVM tries to insert [0, U32_MAX], which overlaps the
    reserved Arm Architecture Calls ranges. KVM_ARM_VM_SMCCC_FILTER then
    returns -EEXIST instead of the -EINVAL that the smccc_filter selftest
    expects for an empty range.
    
    Reject a zero function count explicitly.
    
    Tested with a userspace reproducer on an arm64 VHE host under QEMU TCG:
    EEXIST before, EINVAL after.
    
    Fixes: 821d935c87bc ("KVM: arm64: Introduce support for userspace SMCCC filtering")
    Assisted-by: LLM
    Signed-off-by: Karl Mehltretter <[email protected]>
    Reviewed-by: Steffen Eiden <[email protected]>
    Reviewed-by: Fuad Tabba <[email protected]>
    Tested-by: Fuad Tabba <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: arm64: Transfer the hyp stack pages out of the host stage-2 [+ + +]
Author: Fuad Tabba <[email protected]>
Date:   Tue Sep 8 12:07:10 2026 +0100

    KVM: arm64: Transfer the hyp stack pages out of the host stage-2
    
    commit 3a8c562892b96f35bba1e00d5e455a15963bbb92 upstream.
    
    fix_host_ownership() walks only the linear-map alias of each memblock
    region, and the per-CPU hyp stack, mapped in the private VA range for
    its guard page, has none.
    
    Walk each stack's VA range with the same walker.
    
    Fixes: 1a919b17ef012 ("KVM: arm64: Add guard pages for pKVM (protected nVHE) hypervisor stack")
    Reported-by: Hiroyuki Katsura <[email protected]>
    Cc: [email protected]
    Signed-off-by: Fuad Tabba <[email protected]>
    Reviewed-by: Vincent Donnefort <[email protected]>
    Tested-by: Vincent Donnefort <[email protected]>
    Reviewed-by: Marc Zyngier <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: arm64: Validate the SVE vector length in pkvm_vcpu_init_sve() [+ + +]
Author: Fuad Tabba <[email protected]>
Date:   Tue Aug 25 09:59:45 2026 +0100

    KVM: arm64: Validate the SVE vector length in pkvm_vcpu_init_sve()
    
    [ Upstream commit 2a2eb10795a1e495aebc7f829ccecb72c05b4fd9 ]
    
    pkvm_vcpu_init_sve() clamps only the upper bound of the host-provided
    sve_max_vl, so an invalid vector length reaches sve_state_size_from_vl()
    and the WARN_ON() there, which is fatal at EL2. The existing
    !sve_state_size test rejects such a length, but only after the macro has
    run.
    
    Check sve_vl_valid() before deriving the state size. A valid length
    cannot yield a zero size, so the !sve_state_size test goes with it.
    
    Fixes: 5db1bef93342 ("KVM: arm64: Track SVE state in the hypervisor vcpu structure")
    Reported-by: Stefan Teodorescu <[email protected]>
    Reviewed-by: Marc Zyngier <[email protected]>
    Signed-off-by: Fuad Tabba <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: arm64: vgic-its: Free the caches when GITS_BASER changes [+ + +]
Author: Fuad Tabba <[email protected]>
Date:   Fri Aug 21 07:44:42 2026 +0100

    KVM: arm64: vgic-its: Free the caches when GITS_BASER changes
    
    [ Upstream commit 8cd92f77ae4f5371a7d581f8324c24919670b304 ]
    
    A guest that disables the ITS and re-points or shrinks GITS_BASER<n>
    with VALID still set keeps the devices and collections it mapped
    against the old table, as KVM frees them only when VALID is cleared.
    The contents of the table are IMPLEMENTATION DEFINED, so a write that
    gives GITS_BASER<n> a different address or size may lose whatever the
    old value described. Free the list whenever the stored value changes,
    and drop the translation cache with it.
    
    The cache is not empty just because the ITS is disabled: its->enabled
    is written under the cmd_lock, while vgic_its_resolve_lpi() tests it
    under the its_lock, so an injection can still cache an entry after the
    ITS was disabled. Hence the invalidation inside the its_lock section.
    
    Test for a change rather than a write: its_restore_enable() rewrites
    GITS_BASER<n> from its probe-time cache on resume, and KVM reports
    GITS_TYPER.HCC as 0, so nothing re-maps the boot CPU's collection
    afterwards.
    
    Fixes: 36d6961c2b481 ("KVM: arm/arm64: vgic-its: Free caches when GITS_BASER Valid bit is cleared")
    Suggested-by: Marc Zyngier <[email protected]>
    Link: https://lore.kernel.org/all/[email protected]/
    Signed-off-by: Fuad Tabba <[email protected]>
    Reviewed-by: Marc Zyngier <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: arm64: vgic-its: Skip unreachable devices instead of failing the save [+ + +]
Author: Fuad Tabba <[email protected]>
Date:   Fri Aug 21 07:44:44 2026 +0100

    KVM: arm64: vgic-its: Skip unreachable devices instead of failing the save
    
    [ Upstream commit cc5d96036e01ac330d24b2f0c336d60f82ab4930 ]
    
    vgic_its_save_device_tables() aborts with -EINVAL when a device's entry
    falls outside the device table, which a guest can arrange on its own: an
    indirect table lets it clear an L1 entry's valid bit without touching
    GITS_BASER. That fails a save userspace should be able to issue
    reliably.
    
    Skip the device instead, and point the saved DTE chain past it, as
    commit ad1e686e2378d ("KVM: arm64: vgic-its: Point saved ITEs at the
    next valid entry") does for ITEs. compute_next_devid_offset() takes the
    next device off the list whether or not it was saved, so the predecessor
    would otherwise point at an entry the save never wrote. Restore follows
    that offset while it stays inside the table being scanned: within an L2
    block, or anywhere in a flat table. Both need userspace to remove a
    memslot under the table, since dropping an L1 entry takes the whole
    block with it and scan_its_table() stops at the block boundary.
    
    Fixes: 57a9a117154c9 ("KVM: arm64: vgic-its: Device table save/restore")
    Suggested-by: Marc Zyngier <[email protected]>
    Link: https://lore.kernel.org/all/[email protected]/
    Signed-off-by: Fuad Tabba <[email protected]>
    Reviewed-by: Marc Zyngier <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: Don't treat reserved xarray entries as having memory attributes [+ + +]
Author: Zeng Chi <[email protected]>
Date:   Mon Sep 21 18:24:42 2026 +0800

    KVM: Don't treat reserved xarray entries as having memory attributes
    
    commit 277d3623d99a4fc2623bfb7d191b649ca380605d upstream.
    
    kvm_vm_set_mem_attributes() reserves an xarray entry for every gfn in
    the range before storing the new attributes, so that the store loop
    can't fail partway through.  If one of the reservations fails, e.g. with
    -ENOMEM, the entries that were already reserved are left in the array.
    That is harmless as far as xa_reserve() is concerned, as the reserved
    entries read back as NULL via xa_load(), but it confuses the "does this
    range have no attributes at all" check:
    
            if (!attrs)
                    return !xas_find(&xas, end - 1);
    
    A reserved entry is XA_ZERO_ENTRY, not NULL, and xas_find() returns it
    as present.  So a leftover reservation makes KVM report that a fully
    shared range has attributes even though kvm_get_memory_attributes()
    returns none for every gfn in the range.  On x86, the next time
    mixed-attribute tracking is recomputed for the range (memslot creation,
    or a later attribute change that straddles the 2MiB page),
    hugepage_has_attrs() treats a fully shared 2MiB range as mixed and
    refuses to map it with a hugepage, until userspace happens to set
    attributes on the range again.
    
    Drop the shortcut and handle the !attrs case in the per-index loop,
    using xas_next_entry() to find the next non-NULL entry.  xas_next_entry()
    is essentially an optimized xas_find(), so the effective change is that
    the !attrs lookup now goes through xas_retry() like the attrs != 0 case,
    i.e. reserved entries are skipped and retry entries restart the walk.
    Don't check the index when no entry is found, as the xarray leaves the
    xas index in a bogus state in that case; no entry simply means the rest
    of the range has no attributes.
    
    KVM never stores a non-NULL entry with a value of zero (clearing stores
    NULL), but such an entry would be returned by xas_next_entry() and trip
    the index check, so WARN if one is ever seen.
    
    Fixes: 5a475554db1e ("KVM: Introduce per-page memory attributes")
    Cc: [email protected]
    Suggested-by: Sean Christopherson <[email protected]>
    Cc: David Ballesteros <[email protected]>
    Signed-off-by: Zeng Chi <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    [sean: expand comment to elaborate on xarray APIs, split optimization out]
    Signed-off-by: Sean Christopherson <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg [+ + +]
Author: David Ballesteros <[email protected]>
Date:   Tue Sep 15 17:53:57 2026 +0000

    KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg
    
    commit 382e5d514b6f35bdda2ab9044b4eed23d2ec4254 upstream.
    
    Explicitly instantiate the memory attributes xarray with XA_FLAGS_ACCOUNT
    to ensure that all allocations are accounted to the memcg.  Frustratingly,
    memory allocations done in the "fastpath" do not honor the passed in gfp,
    even for an explicit xa_reserve().  Only the rare, slow path __xas_nomem()
    honors the original gfp.  E.g.
    
      xa_reserve(..., GFP_KERNEL_ACCOUNT)
      |
      -> ...
         |
         -> __xa_cmpxchg_raw()
            |
            -> xas_store()  <== does not take @gfp
               |
               -> xas_create()
                  |
                  -> xas_alloc()
    
    The bug was confirmed by observing that a process in a cgroup limited to
    256 MiB grew radix_tree_node slab by ~512 MiB while its memory.current
    stayed near 0.
    
    Fixes: 5a475554db1e ("KVM: Introduce per-page memory attributes")
    Cc: [email protected]
    Assisted-by: Claude-Code:claude-opus-5
    Signed-off-by: David Ballesteros <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    [sean: rewrite changelog, tag for stable]
    Signed-off-by: Sean Christopherson <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: move TSS constants from kvm_host.h to tss.h [+ + +]
Author: Paolo Bonzini <[email protected]>
Date:   Wed Sep 30 09:12:52 2026 -0400

    KVM: move TSS constants from kvm_host.h to tss.h
    
    [ Upstream commit 35fdfa632e7cf8271189c7ae6f97e7c4b55f7f80 ]
    
    Suggested-by: Kai Huang <[email protected]>
    Signed-off-by: Paolo Bonzini <[email protected]>
    Stable-dep-of: 10180a277549 ("KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails")
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: riscv: Fix NACL hfence entry update order [+ + +]
Author: Zongmin Zhou <[email protected]>
Date:   Wed Aug 26 15:50:09 2026 +0800

    KVM: riscv: Fix NACL hfence entry update order
    
    [ Upstream commit b3d346838ec65fac7fd83f5dbcedd13cadfffddb ]
    
    The SBI v3.0 specification (section 15.1.2) requires a nested HFENCE
    entry to be populated as follows:
    1) find an unused entry with Config.Pending == 0
    2) update the Page_Number and Page_Count words
    3) update the Config word with Config.Pending set
    
    __kvm_riscv_nacl_hfence() writes the Config word first, so the SBI
    implementation (or NACL hardware) can observe a pending entry with
    pnum/pcount values left over from the previous use of that entry,
    resulting in incorrect TLB flush ranges.
    
    Write pnum and pcount first and the Config word last. Since the
    consumer is an external agent on coherent shared memory, use
    WRITE_ONCE() to stop the compiler from reordering the stores and
    smp_wmb() to make the parameter words globally visible before the
    Pending bit is set.
    
    Fixes: d466c19cead5 ("RISC-V: KVM: Add common nested acceleration support")
    Signed-off-by: Zongmin Zhou <[email protected]>
    Reviewed-by: Anup Patel <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Anup Patel <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: s390: Add missing srcu in kvm_s390_set_irq_state() [+ + +]
Author: Claudio Imbrenda <[email protected]>
Date:   Fri Aug 28 13:54:37 2026 +0200

    KVM: s390: Add missing srcu in kvm_s390_set_irq_state()
    
    [ Upstream commit f3a557067d57ce6ae98d485c8009b223e16f5f36 ]
    
    Like kvm_s390_inject_vcpu(), kvm_s390_set_irq_state() also needs the
    kvm->srcu or the slots lock when performing the Store status operation.
    
    Fix by taking kvm->srcu in kvm_s390_set_irq_state().
    
    Fixes: ba5c1e9b6cee ("KVM: s390: interrupt subsystem, cpu timer, waitpsw")
    Fixes: 062e44a9319f ("KVM: s390: Use srcu in kvm_arch_vcpu_unlocked_ioctl()")
    Signed-off-by: Claudio Imbrenda <[email protected]>
    Message-ID: <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: s390: Fix _gaccess_shadow_fault() [+ + +]
Author: Claudio Imbrenda <[email protected]>
Date:   Fri Aug 28 13:54:34 2026 +0200

    KVM: s390: Fix _gaccess_shadow_fault()
    
    [ Upstream commit faff4c8ff3dbec6d71b89dd281a0d17c4478fa44 ]
    
    In some circumstances, it is possible that the page of nested guest
    memory that is being shadowed is not present at all in the parent guest
    gmap. dat_entry_walk() will not find any leaf entry and return with
    -ENOENT, which will erroneously be propagated all the way to userspace.
    
    Fix by manually calling gmap_link() on the memory of the nested guest
    that is being shadowed if the mapping was not already present.
    
    Fixes: e38c884df921 ("KVM: s390: Switch to new gmap")
    Signed-off-by: Claudio Imbrenda <[email protected]>
    Message-ID: <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: s390: Fix dirty marking in adapter_indicators_set*() [+ + +]
Author: Claudio Imbrenda <[email protected]>
Date:   Fri Aug 28 13:54:32 2026 +0200

    KVM: s390: Fix dirty marking in adapter_indicators_set*()
    
    [ Upstream commit d215f014b3526a9898d87bfb4b866287f0864cc9 ]
    
    When the indicator and/or summary bits are set in the guest, the
    accessed page was only marked dirty in KVM if the access was performed
    using the slow path; accesses through the new kvm_arch_set_irq_inatomic
    fast inject path would not mark the page as dirty.
    
    Fix by adding/moving the missing calls to mark_page_dirty(). Note that
    for the inatomic path set_page_dirty{,_lock}() is not needed as the
    page stays pinned; the unpin path correctly marks it as dirty.
    
    Opportunistically reorder the local variables to be in reverse
    Christmas tree order and refactor to use guard().
    
    Fixes: 1e95e3bc6b05 ("KVM: s390: Enable adapter_indicators_set to use mapped pages")
    Signed-off-by: Claudio Imbrenda <[email protected]>
    Message-ID: <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: s390: Fix IRQ injection with SIGP Stop and Store Status [+ + +]
Author: Claudio Imbrenda <[email protected]>
Date:   Wed Aug 12 12:44:33 2026 +0200

    KVM: s390: Fix IRQ injection with SIGP Stop and Store Status
    
    [ Upstream commit d343407b728a80b74be3c24b59f15e60289ea527 ]
    
    When __inject_sigp_stop() is called for a Stop and Store Status
    operation, if the vCPU is running, the interrupt is marked as pending
    and the status is stored by the thread performing the KVM_RUN IOCTL.
    
    If the vCPU is already stopped, the status is stored immediately.
    
    Storing the status means writing into userspace, which might fault, and
    __inject_sigp_stop() is called from do_inject_vcpu() which in turn is
    always called holding a spinlock, which is obviously an issue.
    
    Fix this by returning -EWOULDBLOCK from __inject_sigp_stop(), and
    adding a bool flag to indicate whether a store status is needed. The
    callers of do_inject_vcpu() are modified to pass the pointer to the
    bool flag; whenever a Store Status operation is needed, the callers can
    now perform it outside the spinlock.
    
    Opportunistically refactor kvm_s390_set_irq_state() to use
    scoped_guard() and __free().
    
    Fixes: 6cddd432e3da ("KVM: s390: handle stop irqs without action_bits")
    Signed-off-by: Claudio Imbrenda <[email protected]>
    [ Added Fixes tag while picking -- Claudio ]
    Message-ID: <[email protected]>
    Stable-dep-of: f3a557067d57 ("KVM: s390: Add missing srcu in kvm_s390_set_irq_state()")
    Signed-off-by: Sasha Levin <[email protected]>

KVM: s390: Fix potential races in dat skey functions [+ + +]
Author: Claudio Imbrenda <[email protected]>
Date:   Fri Aug 28 13:54:38 2026 +0200

    KVM: s390: Fix potential races in dat skey functions
    
    [ Upstream commit 27554b9505ddfc0aeab466aeb60929dfa17284c7 ]
    
    When dat_cond_set_storage_key() finds a large page, it will
    conditionally set the storage key in absolute memory using
    large_crste_to_phys() to get the absolute address.
    
    There is a race window between dat_entry_walk() and
    large_crste_to_phys(): the large page could have been split
    concurrently, and large_crste_to_phys() might be called with a crste
    that does not designate a large page, leading to crashes.
    
    Similar issues were also present in dat_set_storage_key().
    
    dat_get_storage_key() and dat_reset_reference_bit() did instead check
    for a potential concurrent splitting of the large page, but then
    handled it incorrectly.
    
    Fix by performing a READ_ONCE on the crste pointer, checking and using
    the result, instead of dereferencing the pointer again. In case a race
    is detacted, try dat_entry_walk() again.
    
    Fixes: 8e03e8316eb2 ("KVM: s390: KVM page table management functions: storage keys")
    Signed-off-by: Claudio Imbrenda <[email protected]>
    Message-ID: <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: s390: Fix race in _destroy_pages_crste() [+ + +]
Author: Claudio Imbrenda <[email protected]>
Date:   Fri Aug 28 13:54:39 2026 +0200

    KVM: s390: Fix race in _destroy_pages_crste()
    
    [ Upstream commit 4ca00a9154f998116fba9a32cce5bd938d228065 ]
    
    Use READ_ONCE() in _destroy_pages_crste() to read the crste, avoid
    dereferencing the pointer multiple times.
    
    Fixes: a2c17f9270cc ("KVM: s390: New gmap code")
    Signed-off-by: Claudio Imbrenda <[email protected]>
    Message-ID: <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: s390: Properly handle NULL pointer in dat_cond_set_storage_key() [+ + +]
Author: Claudio Imbrenda <[email protected]>
Date:   Wed Aug 12 12:44:28 2026 +0200

    KVM: s390: Properly handle NULL pointer in dat_cond_set_storage_key()
    
    [ Upstream commit 88e22ffd1e46b95e40a6afabe486deb1d31a3ae1 ]
    
    Some callers pass NULL as oldkey. Calling page_cond_set_storage_key()
    will cause that NULL pointer to get dereferenced.
    
    Fix by checking for NULL and assigning the pointer to a dummy local
    variable to avoid crashes.
    
    Fixes: 8e03e8316eb2 ("KVM: s390: KVM page table management functions: storage keys")
    Reviewed-by: Christian Borntraeger <[email protected]>
    Reviewed-by: Christoph Schlameuss <[email protected]>
    Signed-off-by: Claudio Imbrenda <[email protected]>
    Message-ID: <[email protected]>
    Stable-dep-of: 27554b9505dd ("KVM: s390: Fix potential races in dat skey functions")
    Signed-off-by: Sasha Levin <[email protected]>

KVM: selftests: fix steal_time for arm64 with host page size > 4K [+ + +]
Author: Sebastian Ott <[email protected]>
Date:   Mon Sep 14 15:10:13 2026 +0200

    KVM: selftests: fix steal_time for arm64 with host page size > 4K
    
    [ Upstream commit 96e6757cb0674acb86ee8b558ee6bffe0eec0bb0 ]
    
    Fix the following failure when running with 16K host page size:
    ==== Test Assertion Failure ====
      lib/kvm_util.c:991: vm_adjust_num_guest_pages(vm->mode, npages) == npages
      pid=873 tid=873 errno=0 - Success
         1  0x0000000000405a27: vm_mem_add at kvm_util.c:991
         2  0x000000000040241f: check_steal_time_uapi at steal_time.c:223 (discriminator 7)
         3   (inlined by) main at steal_time.c:539 (discriminator 7)
         4  0x00007fff8b57af3b: ?? ??:0
         5  0x00007fff8b57b007: ?? ??:0
         6  0x0000000000402b6f: _start at ??:?
      Number of guest pages is not compatible with the host. Try npages=4
    
    Fixes: fc240715fc50 ("KVM: selftests: arm64: Fix steal_time test after UAPI refactoring")
    Reported-by: Zenghui Yu <[email protected]>
    Link: https://lore.kernel.org/kvmarm/[email protected]/T/#u
    Signed-off-by: Sebastian Ott <[email protected]>
    Reviewed-by: Zenghui Yu (Huawei) <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Oliver Upton <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: SEV: Do cache maintenance on the source VM during intra-host migration [+ + +]
Author: Sean Christopherson <[email protected]>
Date:   Wed Sep 23 09:37:21 2026 -0700

    KVM: SEV: Do cache maintenance on the source VM during intra-host migration
    
    commit 93de2a6a4b91b72607136dd656edf03fb399d27f upstream.
    
    Manually perform cache maintenance on the source VM during intra-host
    migration to ensure no stale data is left in CPU caches after the VM is
    destroyed.  Because the source VM is "converted" to a non-SEV VM, KVM's
    memory reclaim flows won't trigger cache maintenance, e.g. when all guest
    memory is reclaimed in response to detaching from the mmu_notifier.
    
    Note, relying on the destination VM to do cache maintenance isn't an option
    as KVM doesn't require identical guest memory configurations, i.e. the
    source VM may have access to memory that the destination VM does not.
    Enforcing equivalent memory configurations is infeasible, as it would
    require a *deep* comparison of memslots, e.g. to verify that not only are
    the memslot identical, but what the memslots point at is also identical.
    
    Fixes: b56639318bb2 ("KVM: SEV: Add support for SEV intra host migration")
    Cc: [email protected]
    Reported-by: Stefan Teodorescu <[email protected]>
    Signed-off-by: Sean Christopherson <[email protected]>
    Message-ID: <[email protected]>
    Signed-off-by: Paolo Bonzini <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV [+ + +]
Author: Sean Christopherson <[email protected]>
Date:   Wed Sep 23 09:37:20 2026 -0700

    KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV
    
    commit 12c1f6e03f944e399bd2c88441dca5dc702b95a5 upstream.
    
    Unconditionally free SEV's "have run CPUs" cpumask in the VM destroy path,
    i.e. even for what appear to be non-SEV VMs, as an SEV VM becomes a non-SEV
    VM if its state is intra-host migrated.  Alternatively, the mask could be
    freed in sev_migrate_from() when "converting" the source VM, but that gets
    annoying because ideally KVM would nullify the mask to guard against UAF,
    and nullifying the mask would need be conditioned on CPUMASK_OFFSTACK=y.
    
    Freeing the mask during sev_migrate_from() is also not robust against other
    KVM bugs, though that's kind of a moot point since any such bugs would show
    up even if sev->active is never set.  I.e. KVM must get that side of things
    correct.  But, that's not a great reason to add more code just to make
    things marginally less robust.
    
    Fixes: 6f38f8c57464 ("KVM: SVM: Flush cache only on CPUs running SEV guest")
    Cc: [email protected]
    Reported-by: Stefan Teodorescu <[email protected]>
    Signed-off-by: Sean Christopherson <[email protected]>
    Message-ID: <[email protected]>
    Signed-off-by: Paolo Bonzini <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: x86/mmu: Move kvm_arch_async_page_ready() below kvm_tdp_page_fault() [+ + +]
Author: Sean Christopherson <[email protected]>
Date:   Wed Sep 30 09:12:50 2026 -0400

    KVM: x86/mmu: Move kvm_arch_async_page_ready() below kvm_tdp_page_fault()
    
    [ Upstream commit 31a2cf735cc00c01526bcb676f111d5c0b711d17 ]
    
    Move the implementation of kvm_arch_async_page_ready() "down" in mmu.c so
    that it lives below kvm_tdp_page_fault().  This will allow moving
    kvm_mmu_do_page_fault() into mmu.c without needing a forward declaration.
    
    No functional change intended.
    
    Reviewed-by: Yosry Ahmed <[email protected]>
    Reviewed-by: Kai Huang <[email protected]>
    Signed-off-by: Sean Christopherson <[email protected]>
    Reviewed-by: Binbin Wu <[email protected]>
    Message-ID: <[email protected]>
    Signed-off-by: Paolo Bonzini <[email protected]>
    Stable-dep-of: 10180a277549 ("KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails")
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: x86/mmu: Move kvm_mmu_do_page_fault() from mmu_internal.h => mmu.c [+ + +]
Author: Sean Christopherson <[email protected]>
Date:   Wed Sep 30 09:12:51 2026 -0400

    KVM: x86/mmu: Move kvm_mmu_do_page_fault() from mmu_internal.h => mmu.c
    
    [ Upstream commit 20fe9252460524f9028b515f0f672442ffe6779b ]
    
    Move kvm_mmu_do_page_fault() into mmu.c, as there are no users outside of
    mmu.c, and the function typically isn't inlined by the compiler anyways.
    This will allow moving the EMULTYPE_xxx definitions into x86.h without
    having to include x86.h in mmu_internal.h, i.e. will help preserve the
    goal of making x86.h KVM x86's "top-level" include.
    
    No functional change intended.
    
    Reviewed-by: Yosry Ahmed <[email protected]>
    Reviewed-by: Kai Huang <[email protected]>
    Signed-off-by: Sean Christopherson <[email protected]>
    Reviewed-by: Binbin Wu <[email protected]>
    Message-ID: <[email protected]>
    Signed-off-by: Paolo Bonzini <[email protected]>
    Stable-dep-of: 10180a277549 ("KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails")
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: x86/pmu: Move Intel PMU global MSRs to intel_is_valid_msr() [+ + +]
Author: Jim Mattson <[email protected]>
Date:   Wed Sep 2 11:47:11 2026 -0700

    KVM: x86/pmu: Move Intel PMU global MSRs to intel_is_valid_msr()
    
    [ Upstream commit 79a71cc2568f4b5d42284da2aa26f3b4f47ce01b ]
    
    Commit c85cdc1cc1ea ("KVM: x86/pmu: Move handling PERF_GLOBAL_CTRL and
    friends to common x86") moved the existence check for the following Intel
    PMU MSRs to kvm_pmu_is_valid_msr():
     - MSR_CORE_PERF_GLOBAL_STATUS
     - MSR_CORE_PERF_GLOBAL_CTRL
     - MSR_CORE_PERF_GLOBAL_OVF_CTRL
    
    That commit deemed these MSRs valid whenever pmu->version > 1. It intended
    to share the check with AMD PerfMonV2 because both vendor implementations
    require version 2 or greater for global PMU controls.  However, as noted in
    the commit message, AMD uses different MSR indices for its global PMU
    registers.
    
    Commit 4a2771895ca6 ("KVM: x86/svm/pmu: Add AMD PerfMonV2 support")
    subsequently added AMD PerfMonV2 support and set pmu->version = 2.  Because
    kvm_pmu_is_valid_msr() validated the Intel MSRs whenever pmu->version > 1,
    KVM incorrectly permitted AMD guests with PerfMonV2 to access these Intel
    MSRs without a #GP.
    
    Move the validation of these Intel MSRs to intel_is_valid_msr() and remove
    the common switch statement from kvm_pmu_is_valid_msr(). AMD already
    validates its own global PMU MSRs in amd_is_valid_msr().
    
    Fixes: 4a2771895ca6 ("KVM: x86/svm/pmu: Add AMD PerfMonV2 support")
    Signed-off-by: Jim Mattson <[email protected]>
    Reviewed-by: Like Xu <[email protected]>
    Reviewed-by: Sandipan Das <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sean Christopherson <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

KVM: x86: Add static calls for nested virtualization ops [+ + +]
Author: Sean Christopherson <[email protected]>
Date:   Wed Sep 30 09:12:54 2026 -0400

    KVM: x86: Add static calls for nested virtualization ops
    
    [ Upstream commit 4b9819a50674dd40d176033f9ea2e23a70889dc6 ]
    
    Use static calls to invoke nested virtualization ops, as many of the calls
    are in relatively hot paths when L2 is active, e.g. checking for events,
    and because there's no reason not use static calls these days.
    
    Opportunistically use a RET0 static call for get_evmcs_version() instead
    of manually checking for a non-NULL vendor hook.
    
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sean Christopherson <[email protected]>
    
    Backport notes:
    Preserve the stable tree's kvm_x86_vendor_init()/exit() declarations and
    ignore_msrs/report_ignored_msrs module parameters at the conflicting
    insertion points. Initialize nested static calls in the existing
    kvm_ops_update() instead of adding kvm_nested_ops_update(), preserving the
    same mandatory, optional, and return-zero hook behavior without adding a
    function. Keep the get_nested_state_pages call conversion so target
    10180a277549339020b08000206092c07e0bff5a applies unchanged.
    
    Stable-dep-of: 10180a277549 ("KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails")
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: x86: Move IRQ-related helper declarations from kvm_host.h => irq.h [+ + +]
Author: Sean Christopherson <[email protected]>
Date:   Wed Sep 30 09:12:48 2026 -0400

    KVM: x86: Move IRQ-related helper declarations from kvm_host.h => irq.h
    
    [ Upstream commit 0bdd2d6d7328a7debcf138a7d6ef95d01a26b1b8 ]
    
    Move the function declaration for APIs to get/query pending IRQs from
    kvm_host.h to irq.h, as the APIs are only used by KVM x86 code.
    
    No functional change intended.
    
    Reviewed-by: Yosry Ahmed <[email protected]>
    Signed-off-by: Sean Christopherson <[email protected]>
    Reviewed-by: Kai Huang <[email protected]>
    Reviewed-by: Binbin Wu <[email protected]>
    Message-ID: <[email protected]>
    Signed-off-by: Paolo Bonzini <[email protected]>
    Stable-dep-of: 10180a277549 ("KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails")
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

KVM: x86: Reject nested CAP enablement if nested virtualization is disabled [+ + +]
Author: Sean Christopherson <[email protected]>
Date:   Wed Sep 30 09:12:53 2026 -0400

    KVM: x86: Reject nested CAP enablement if nested virtualization is disabled
    
    [ Upstream commit b65699be2c606d2687593516d93e66d16a713b61 ]
    
    Add a flag to explicitly track if nested virtualization is enabled, and use
    it enumerate that various nested CAPs are unsupported, and to reject
    enablement of said CAPs.  When the nested ops hooks were moved to their
    own structure, KVM's NULL-by-default behavior was deliberately dropped,
    with the changelog asserting that all was well.  That wasn't quite true;
    there is no danger to KVM, but now KVM is over-reporting support for
    KVM_CAP_NESTED_STATE and KVM_CAP_HYPERV_ENLIGHTENED_VMCS.
    
    Fixes: 33b22172452f ("KVM: x86: move nested-related kvm_x86_ops to a separate struct")
    Reviewed-by: Vitaly Kuznetsov <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sean Christopherson <[email protected]>
    Stable-dep-of: 10180a277549 ("KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails")
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
landlock: Work around gcc-16 -Wuninitialized warning [+ + +]
Author: Arnd Bergmann <[email protected]>
Date:   Tue Sep 15 22:10:30 2026 +0200

    landlock: Work around gcc-16 -Wuninitialized warning
    
    commit c4e941bb7654bcbdfb0b6f3341dc2acfdf235c8d upstream.
    
    gcc has a bug with -ftrivial-auto-var-init=pattern that produces a
    warning for correct code that uses sparse bitfields:
    
    security/landlock/fs.c: In function 'is_access_to_paths_allowed.isra':
    security/landlock/fs.c:767:28: error: '_layer_masks_child1' is used uninitialized [-Werror=uninitialized]
      767 |         struct layer_masks _layer_masks_child1, _layer_masks_child2;
          |                            ^~~~~~~~~~~~~~~~~~~
    security/landlock/fs.c:767:28: note: '_layer_masks_child1' declared here
      767 |         struct layer_masks _layer_masks_child1, _layer_masks_child2;
          |                            ^~~~~~~~~~~~~~~~~~~
    security/landlock/fs.c: In function 'hook_unix_find':
    security/landlock/fs.c:1649:28: error: 'layer_masks' is used uninitialized [-Werror=uninitialized]
     1649 |         struct layer_masks layer_masks;
          |                            ^~~~~~~~~~~
    security/landlock/fs.c:1649:28: note: 'layer_masks' declared here
     1649 |         struct layer_masks layer_masks;
          |                            ^~~~~~~~~~~
    
    To work around this, change the definition of struct layer_mask to
    use an explictit padding field.
    
    Link: https://gcc.gnu.org/bugzilla/show_bug.cgi?id=110743
    Link: https://lore.kernel.org/all/[email protected]/
    Fixes: a260c0055665 ("landlock: Add a place for flags to layer rules")
    Signed-off-by: Arnd Bergmann <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    [mic: Use BITS_PER_TYPE(), fix kdoc warnings, fix commit message
    according to v2 changes]
    Cc: [email protected]
    [mic: Backport: use CONFIG_AUDIT to calculate the padding width]
    Signed-off-by: Mickaël Salaün <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
libbpf: Reject truncated ldimm64 CO-RE relocations [+ + +]
Author: Kumar Kartikeya Dwivedi <[email protected]>
Date:   Fri Sep 18 01:32:18 2026 +0200

    libbpf: Reject truncated ldimm64 CO-RE relocations
    
    [ Upstream commit b4e875d397da451fb4e9c573ff4b86db53caba05 ]
    
    CO-RE relocation of an ldimm64 instruction operates on two instruction
    slots. A malformed BPF ELF can end a function after the first slot and
    attach a CO-RE relocation to it. libbpf allocates the instruction array
    according to the function symbol size, so the shared relocation code would
    then access beyond the allocation.
    
    Reject a terminal ldimm64 in libbpf's relocation loop, where the program
    length is available, before resolving or applying the relocation. Both
    resolved and unresolved relocations validate the absent second slot, and
    unresolved relocation poisoning would additionally write past the array.
    
    The in-kernel caller is protected by the verifier's early instruction-stream
    check before it applies CO-RE relocations.
    
    Fixes: eacaaed784e2 ("libbpf: Implement enum value-based CO-RE relocations")
    Reported-by: Sashiko <[email protected]>
    Signed-off-by: Kumar Kartikeya Dwivedi <[email protected]>
    Link: https://lore.kernel.org/[email protected]
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Eduard Zingerman <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
Linux: Linux 7.2.9 [+ + +]
Author: Greg Kroah-Hartman <[email protected]>
Date:   Sat Oct 3 12:43:16 2026 +0200

    Linux 7.2.9
    
    Link: https://lore.kernel.org/r/[email protected]
    Tested-by: Ronald Warsow <[email protected]>
    Tested-by: Peter Schneider <[email protected]>
    Tested-by: Benjamin Boortz <[email protected]>
    Tested-by: Salvatore Bonaccorso <[email protected]>
    Tested-by: Pavel Machek (CIP) <[email protected]>
    Tested-by: Ron Economos <[email protected]>
    Tested-by: Takeshi Ogasawara <[email protected]>
    Tested-by: Miguel Ojeda <[email protected]>
    Tested-by: Brett A C Sheffield <[email protected]>
    Tested-by: Barry K. Nathan <[email protected]>
    Tested-by: Justin M. Forbes <[email protected]>
    Tested-by: Florian Fainelli <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
llc: fix skb UAF and leaks on llc_mac_hdr_init() failure [+ + +]
Author: Eric Dumazet <[email protected]>
Date:   Thu Sep 24 08:29:48 2026 +0000

    llc: fix skb UAF and leaks on llc_mac_hdr_init() failure
    
    [ Upstream commit 72f9dd522f8d6c5a00be9695c7bb74631eb5069e ]
    
    In llc_conn_ac_resend_i_xxx_x_set_0_or_send_rr(), if llc_mac_hdr_init()
    fails, kfree_skb(skb) is called instead of kfree_skb(nskb). This leaks
    the newly allocated nskb, reads from the freed skb via LLC_I_GET_NR(pdu),
    and double-frees skb when llc_conn_state_process() drops its reference.
    
    In llc_sap_action_send_xid_r() and llc_sap_action_send_test_r(), nskb is
    leaked if llc_mac_hdr_init() returns an error.
    
    Free nskb in all three error paths.
    
    Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
    Closes: https://lore.kernel.org/netdev/[email protected]/
    Signed-off-by: Eric Dumazet <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

llc: reserve device headroom for allocated frames [+ + +]
Author: Zixuan Chai <[email protected]>
Date:   Thu Sep 24 09:26:05 2026 +0800

    llc: reserve device headroom for allocated frames
    
    commit 72b5b9a28b996e09b8b5b944370c79851bb68f52 upstream.
    
    llc_alloc_frame() reserves link-layer headroom using the device type.
    This is insufficient for stacked Ethernet devices such as VLAN devices,
    where vlan_dev_hard_header() pushes a VLAN header before the lower
    device's Ethernet header. An LLC response on such a device can
    therefore underflow skb headroom in eth_header().
    
    Use LL_RESERVED_SPACE() to account for the device's actual required
    headroom while preserving the existing LLC device-type check.
    
    Fixes: bf9ae5386bca ("llc: use dev_hard_header")
    Cc: [email protected]
    Reported-by: VEGA <[email protected]>
    Signed-off-by: Zixuan Chai <[email protected]>
    Signed-off-by: Ren Wei <[email protected]>
    Reviewed-by: Eric Dumazet <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
macsec: initialize SecY before registering the netdevice [+ + +]
Author: Haseeb Malik <[email protected]>
Date:   Mon Sep 21 16:40:30 2026 -0400

    macsec: initialize SecY before registering the netdevice
    
    [ Upstream commit c2de369c5c5b8599ca10fd5ca8d11fcd845c1331 ]
    
    Creating a MACsec device with MAC offload over an LRO-capable lower
    device triggers a warning in rtmsg_ifinfo_build_skb() when IPv4
    forwarding is enabled by default.
    
    register_netdevice() invokes inetdev_init(), which disables LRO and emits
    a NETDEV_FEAT_CHANGE notification. This reaches macsec_fill_info() before
    macsec_add_dev() initializes the SecY. key_len is still zero, so
    macsec_fill_info() returns -EMSGSIZE and trips the WARN_ON in
    rtmsg_ifinfo_build_skb(), even though the skb has enough space.
    
    Even without the warning, notifications during registration can report
    uninitialized SecY attributes, including the SCI. This ordering has existed
    since the driver was introduced.
    
    Initialize the SecY and apply the new-link attributes before registration.
    Move MAC address inheritance into macsec_newlink() so the SCI can also be
    initialized before registration-time notifications report it. Move the
    per-CPU statistics and metadata destination allocation into ndo_init(),
    and release partial allocations on failure.
    
    Fixes: c09440f7dcb3 ("macsec: introduce IEEE 802.1AE driver")
    Reported-by: [email protected]
    Closes: https://syzkaller.appspot.com/bug?extid=f2f6312ad1b5a0bfe316
    Suggested-by: Sabrina Dubroca <[email protected]>
    Link: https://lists.openwall.net/linux-kernel/2026/08/19/552
    Signed-off-by: Haseeb Malik <[email protected]>
    Reviewed-by: Sabrina Dubroca <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
mctp: route: iterate socket tag list in mctp_lookup_prealloc_tag() [+ + +]
Author: Hui Peng <[email protected]>
Date:   Mon Sep 21 05:10:01 2026 +0000

    mctp: route: iterate socket tag list in mctp_lookup_prealloc_tag()
    
    commit 26cc0e69cce062cd3aa6fae33074684669c35a71 upstream.
    
    When a socket transmits a packet with MCTP_TAG_PREALLOC set,
    mctp_lookup_prealloc_tag() iterates over the per-netns &mns->keys list
    and matches netid, req_tag, peer_addr, and manual_alloc, without
    checking whether tmp->sk == &msk->sk. This allows any MCTP socket in the
    same network namespace to use and consume another socket's preallocated
    tag.
    
    Iterate the socket's own tag list (&msk->keys via sklist) instead of the
    namespace-wide &mns->keys list in mctp_lookup_prealloc_tag(), ensuring
    that only tags allocated by msk are matched.
    
    Tested in QEMU against Linux 7.3.0-rc3 by allocating a manual tag
    (0x18) on socket A via SIOCMCTPALLOCTAG for peer EID 9 and sending a
    4-byte message with MCTP_TAG_PREALLOC from socket B in the same network
    namespace. On the unfixed kernel, sendto(sock_b) using socket A's
    preallocated tag succeeds (ret = 4); with this patch applied,
    sendto(sock_b) fails with -ENOENT (errno = 2) while sendto(sock_a)
    succeeds (ret = 4).
    
    Fixes: 63ed1aab3d40 ("mctp: Add SIOCMCTP{ALLOC,DROP}TAG ioctls for tag control")
    Suggested-by: Jeremy Kerr <[email protected]>
    Cc: [email protected]
    Signed-off-by: Hui Peng <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
mm/damon/core: allow esz to be set to zero [+ + +]
Author: Liew Rui Yan <[email protected]>
Date:   Tue Sep 8 06:54:11 2026 -0700

    mm/damon/core: allow esz to be set to zero
    
    commit 90179da203ba8b708c84a12a07cd44be0f346334 upstream.
    
    When the temporal quota goal tuner determines that the goal has been
    achieved (score >= 10000), it sets esz_bp to zero so that the esz becomes
    zero.  However, damos_set_effective_quota() clamps the esz to
    min_region_sz when quota->ms is set.
    
    This is a minor issue, the main problem is that it doesn't match the
    description in the documentation, which state that if the goal has already
    been [over-]achieved, the quota will be set to zero.
    
    Fix this by set quota (esz) as minimum as possible.
    
    Link: https://lore.kernel.org/[email protected]
    Fixes: 8bbde987c2b8 ("mm/damon/core: disallow time-quota setting zero esz")
    Signed-off-by: SJ Park <[email protected]>
    Signed-off-by: Liew Rui Yan <[email protected]>
    Signed-off-by: Andrew Morton <[email protected]>
    Reviewed-by: SJ Park <[email protected]>
    Cc: <[email protected]> # v7.1.x
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

mm/damon/core: fix unconditionally skip last region [+ + +]
Author: Liew Rui Yan <[email protected]>
Date:   Tue Sep 8 06:47:38 2026 -0700

    mm/damon/core: fix unconditionally skip last region
    
    commit b3723b596b548c837a766aae3553c14a7b15af2b upstream.
    
    Once quota set, the charge_{target,addr}_from unconditionally skips and
    resets at the last region of the tracked target, so the last region can be
    skipped even when it has not been processed.
    
    Example:
    
        1. Target has 2 regions: R1 (0-100 bytes) and R2 (100-200 bytes).
        2. Quota is configured to process only 100 bytes per window.
        3. Window 1: Processes R1 (0-100).  Quota is full.  charge_{target,
           addr}_from is saved at (Target, 100).
        4. Window 2: The loop reaches R2.  Because R2 is
           damon_last_region(t), the old code unconditionally returns true,
           skipping R2 entirely and resetting the charge_{target,addr}_from.
    
        Result: R2 is permanently skipped even though it has never been
        processed.
    
    However, it is important to note that this is a very minor issue.  This is
    because it is triggered only when the previous window saved/kept
    charge_{target,addr}_from, and in the next window, all regions except the
    last region were skipped by damos_skip_charged_region().
    
    Fix this by only resetting the charge_{target,addr}_from when last region
    is reached, only skipping when it is applied or cannot split.
    
    Link: https://lore.kernel.org/[email protected]
    Fixes: 50585192bc2e ("mm/damon/schemes: skip already charged targets and regions")
    Signed-off-by: Liew Rui Yan <[email protected]>
    Reviewed-by: SJ Park <[email protected]>
    Signed-off-by: SJ Park <[email protected]>
    Signed-off-by: Andrew Morton <[email protected]>
    Cc: <[email protected]> # v5.16.x
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

mm/damon/core: reset invalid quota->charge_target_from [+ + +]
Author: SJ Park <[email protected]>
Date:   Thu Sep 10 07:28:45 2026 -0700

    mm/damon/core: reset invalid quota->charge_target_from
    
    commit eb64948249781bda35de04feab5a0acc36aa9051 upstream.
    
    DAMOS can suddenly stop working if a target process that the quota is just
    fully charged on is terminated.  Fix by catching and processing the corner
    case.
    
    When DAMOS quota is fully charged, the target and the region to continue
    applying the action in the next round is saved in
    damos_quota->charge_{target,addr}_from.  In the next round, DAMOS iterates
    targets and regions from the beginning.  It skips applying the action to
    the regions until it visits and skips the saved target/region.
    
    Virtual address space targets become invalid if the process is terminated.
    Trying to apply the scheme to invalid target is just a waste of time.
    Hence commit 6e4930e33329 ("mm/damon/core: fix wasteful CPU calls by
    skipping non-existent targets") made the logic to skip invalid targets.
    However, it does skip before the charged target/region skipping/updating.
    
    Let's suppose the user runs DAMOS for multiple virtual address spaces with
    a quota.  The quota exceeded in the middle of a virtual address space.
    And the process of the address space is terminated.  Then the
    charge_target_from points to the invalid target.  The pointer update logic
    is skipped for the invalid target, so the charge_target_from is never
    updated.  DAMOS action to every target/region is skipped.  From the user's
    perspective, it would look like suddenly DAMOS has stopped working.
    
    No critical leak or crash can happen.  The user could reinstall the
    scheme.  But this makes use of DAMOS under certain setups quite
    unreliable.
    
    When the invalid target is found, further check the corner case and reset
    the pointer.
    
    This issue was discovered [1] by Sashiko.
    
    Link: https://lore.kernel.org/[email protected]
    Link: https://lore.kernel.org/[email protected] [1]
    Fixes: 6e4930e33329 ("mm/damon/core: fix wasteful CPU calls by skipping non-existent targets")
    Signed-off-by: SJ Park <[email protected]>
    Signed-off-by: Andrew Morton <[email protected]>
    Cc: Enze Li <[email protected]>
    Cc: <[email protected]> # 7.0.x
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
mm/damon/ops-common: use a page-aligned address in damon_ptep_mkold() [+ + +]
Author: Nathan Gao <[email protected]>
Date:   Thu Sep 3 17:28:27 2026 -0700

    mm/damon/ops-common: use a page-aligned address in damon_ptep_mkold()
    
    commit f166586f74dd5d9cbadaabf86ef81c8ddf6cafa7 upstream.
    
    __damon_va_prepare_access_check() picks a random byte address within the
    region and stores it in r->sampling_addr.  damon_va_mkold() passes it into
    a page table walk, which hands it to damon_ptep_mkold() as the address of
    the page to sample:
    
      damon_va_mkold(mm, r->sampling_addr)
        damon_va_walk_page_range(mm, addr, addr + 1)
          damon_mkold_pmd_entry()
            damon_ptep_mkold(pte, vma, addr)
              ptep_test_and_clear_young(vma, addr, pte)
              mmu_notifier_clear_young(mm, addr, addr + PAGE_SIZE)
    
    For arm64, before commit 6f0e1142173a ("arm64: mm: support batch clearing
    of the young flag for large folios"), the contpte helper walked exactly
    CONT_PTES entries from the aligned-down page table pointer and used @addr
    only to pass down to each entry, so an unaligned value was harmless:
    
            ptep = contpte_align_down(ptep);
            addr = ALIGN_DOWN(addr, CONT_PTE_SIZE);
            for (i = 0; i < CONT_PTES; i++, ptep++, addr += PAGE_SIZE)
    
    Now the range to walk is derived from @addr instead: end = addr + nr *
    PAGE_SIZE, rounded up to CONT_PTE_SIZE.  For a sample in the last page of
    a contpte block, the sub-page offset puts end just past the block
    boundary, so the round-up lands a whole block further and the walk clears
    PTE_AF in CONT_PTES entries beyond the sampled block.  For the last block
    in a page table page, those entries are past the end of that page, so the
    walk writes into the page that follows.
    
    Triggered by the full 7.1/7.2 kernel selftest suite on arm64 (EC2
    c/m6g.4xlarge).  The kernel sometimes crashes at or shortly after the
    DAMON test.
    
    What the overrun does depends on the page that happens to follow the page
    table, so there is no single signature.  If that page is read-only, the
    write faults in the sampling path itself:
    
      Unable to handle kernel write to read-only memory at virtual address ffff0003c5d2d000
        FSC = 0x0f: level 3 permission fault
        CM = 0, WnR = 1, TnD = 0, TagAccess = 0
      CPU: 10 UID: 0 PID: 3487 Comm: kdamond.2
      pc : contpte_test_and_clear_young_ptes+0x70/0xc0
      lr : damon_ptep_mkold+0x1e8/0x1f8
      Call trace:
       contpte_test_and_clear_young_ptes+0x70/0xc0 (P)
       damon_mkold_pmd_entry+0x150/0x170
       walk_pmd_range+0x110/0x2b0
       walk_pud_range+0x10c/0x208
       walk_pgd_range+0x134/0x258
       __walk_page_range+0x98/0x1b0
       walk_page_range_vma_unsafe+0x90/0x148
       walk_page_range_vma+0x28/0x40
       damon_va_walk_page_range+0x114/0x2b8
       damon_va_prepare_access_checks+0xec/0x1a8
       kdamond_fn+0x534/0x770
       kthread+0x128/0x138
       ret_from_fork+0x10/0x20
    
    Otherwise the page is writable, the PTE_AF clearing succeeds silently and
    the damage only surfaces later, in whatever happened to own the page, so
    the backtrace is unrelated to DAMON and differs between runs.
    
    Pass a page-aligned address to the ptep_test_and_clear_young() call in
    damon_ptep_mkold(), which is the only place DAMON can reach
    contpte_test_and_clear_young_ptes() from.  Nothing else sees the aligned
    address, and r->sampling_addr itself is left as is, so the sampling and
    region bookkeeping semantics are unchanged.
    
    Link: https://lore.kernel.org/[email protected]
    Fixes: 6f0e1142173a ("arm64: mm: support batch clearing of the young flag for large folios")
    Signed-off-by: Nathan Gao <[email protected]>
    Signed-off-by: SJ Park <[email protected]>
    Signed-off-by: Andrew Morton <[email protected]>
    Reviewed-by: SJ Park <[email protected]>
    Reviewed-by: Baolin Wang <[email protected]>
    Cc: Baolin Wang <[email protected]>
    Cc: David Hildenbrand (Arm) <[email protected]>
    Cc: Ryan Roberts <[email protected]>
    Cc: <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
mm/damon/vaddr: avoid hw-driven pte updates during damon_hugetlb_mkold() [+ + +]
Author: SJ Park <[email protected]>
Date:   Mon Sep 7 10:03:56 2026 -0700

    mm/damon/vaddr: avoid hw-driven pte updates during damon_hugetlb_mkold()
    
    commit 39c0ceedd54557bdc1542de08d22b2ed33e534e4 upstream.
    
    damon_hugetlb_mkold() reads the page table entry into a local variable,
    unsets the accessed bit in the variable, and updates the page table entry
    with the updated variable value.  If hardware updates the same page table
    entry in parallel, the hw updates could be lost.  For example,
    hardware-updated dirty bits might be lost.
    
    Avoid the parallel updates by clearing the page table entry when reading
    it together, using huge_ptep_get_and_clear().  If a parallel write to the
    memory is made after the clearing, the hw will see the page table entry is
    cleared, trigger page fault and wait until it is handled.  The page fault
    handling will wait for damon_hugetlb_mkold() due to the page table lock.
    
    Because hugetlbfs is an in-memory file system and hugetlb pages cannot be
    reclaimed, no critical issue is expected to my best knowledge.  But
    definitely this is a nasty bug that should be fixed sooner rather than
    later.
    
    The issue was discovered [1] by Sashiko.
    
    Link: https://lore.kernel.org/[email protected]
    Link: https://lore.kernel.org/[email protected] [1]
    Fixes: 49f4203aae06 ("mm/damon: add access checking for hugetlb pages")
    Signed-off-by: SJ Park <[email protected]>
    Signed-off-by: Andrew Morton <[email protected]>
    Cc: Baolin Wang <[email protected]>
    Cc: <[email protected]> # 5.17.x
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
mm/hugetlb: do not dissolve gigantic pages without runtime support [+ + +]
Author: Longlong Xia <[email protected]>
Date:   Sun Aug 23 12:40:51 2026 +0800

    mm/hugetlb: do not dissolve gigantic pages without runtime support
    
    commit a363c62a653cc8b3e21da9545fa4e028ef50f9c3 upstream.
    
    dissolve_free_hugetlb_folio() doesn't check
    hstate_is_gigantic_no_runtime(h) though remove_hugetlb_folio()/
    update_and_free_hugetlb_folio() silently bail for such folios, so it frees
    a still-listed folio and, on vmemmap restore failure, the
    add_hugetlb_folio() rollback corrupts the free list.
    
    Link: https://lore.kernel.org/[email protected]
    Fixes: 6eb4e88a6d27 ("hugetlb: create remove_hugetlb_page() to separate functionality")
    Signed-off-by: Longlong Xia <[email protected]>
    Signed-off-by: Andrew Morton <[email protected]>
    Assisted-by: Codex:gpt-5.6-sol
    Acked-by: Muchun Song <[email protected]>
    Cc: David Hildenbrand <[email protected]>
    Cc: Miaohe Lin <[email protected]>
    Cc: Michal Hocko <[email protected]>
    Cc: Oscar Salvador <[email protected]>
    Cc: <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

mm/hugetlb: preserve mremap address delta when skipping page tables [+ + +]
Author: Jaewook You <[email protected]>
Date:   Mon Sep 14 22:23:52 2026 +0900

    mm/hugetlb: preserve mremap address delta when skipping page tables
    
    commit 9bdad082d44bdcf93716973dcba6be77e8a06e7b upstream.
    
    move_hugetlb_page_tables() optimizes mremap() by advancing to the last
    entry in the page table when the source page table does not exist, either
    initially or after unsharing a PMD table.  The common loop increment then
    steps to the first entry in the next page table.
    
    However, the code advances both the source and destination addresses to
    the last entries in their respective page tables, which is wrong.  The
    destination address must be advanced only by the same amount as the source
    address.
    
    If the source and destination offsets within their page tables differ, the
    destination address can be advanced too far, causing follow-up issues.
    Fix this by advancing the destination address by the source advance
    distance.
    
    With a reproducer, we were able to trigger a kernel panic on x86-64.  With
    this fix in place, we can no longer reproduce the issue.
    
    Link: https://lore.kernel.org/[email protected]
    Fixes: e95a9851787b ("hugetlb: skip to end of PT page mapping when pte not present")
    Fixes: 4ddb4d91b82f ("hugetlb: do not update address in huge_pmd_unshare")
    Signed-off-by: Jaewook You <[email protected]>
    Signed-off-by: Andrew Morton <[email protected]>
    Acked-by: David Hildenbrand (Arm) <[email protected]>
    Cc: Johan Hovold <[email protected]>
    Cc: Muchun Song <[email protected]>
    Cc: Oscar Salvador <[email protected]>
    Cc: <[email protected]>
    Assisted-by: LLM
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
mm/rmap: fix missing barrier between anon_vma init and vma->anon_vma publish [+ + +]
Author: Jinjiang Tu <[email protected]>
Date:   Tue Sep 8 20:29:24 2026 +0800

    mm/rmap: fix missing barrier between anon_vma init and vma->anon_vma publish
    
    commit b6ac0b3f6013c168f22cad97e79967accacb08e1 upstream.
    
    On arm64 server, we find that a task trying to grab the anon_vma lock
    triggers hungtask.
    
    INFO: task main:2354726 blocked for more than 120 seconds.
          Tainted: G            E     5.10.0-0021.aarch64 #1
    "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
    task:main            state:D stack:    0 pid:2354726 ppid:2350673 flags:0x00000a01
    Call trace:
     __switch_to+0x7c/0xbc
     __schedule+0x3b4/0x8a0
     schedule+0x50/0xe0
     rwsem_down_write_slowpath+0x3cc/0x6cc
     down_write+0x60/0x260
     __anon_vma_prepare+0x6c/0x210
     do_anonymous_page+0x258/0x660
     handle_pte_fault+0x188/0x214
     __handle_mm_fault+0x1b0/0x380
     handle_mm_fault+0xf4/0x284
     do_page_fault+0x19c/0x494
     do_translation_fault+0xcc/0xf8
     do_mem_abort+0x48/0xac
     el0_da+0x44/0x80
     el0_sync_handler+0x88/0xb4
     el0_sync+0x160/0x180
    
    After analyzing the vmcore, we found the anon_vma->root->rwsem.count is
    -1.  There is another anon_vma whose anon_vma->root->rwsem.count is 1, the
    anon_vma->root->rwsem.owner shows the lock is held, but the stack of the
    task shows the task doesn't hold the anon_vma lock.
    
    After adding more debugging info, we found __anon_vma_prepare() reuses
    anon_vma and triggers the UAF of anon_vma->root due to missing memory
    barrier, leading to locking and unlocking two different anon_vma->root,
    thus leading to an anon_vma will never be unlocked, and another anon_vma
    couldn't be locked anymore.
    
    This race requires two adjacent VMAs that are not merged but are
    anon_vma-compatible (e.g., they differ in VMA_ACCESS_FLAGS that can be
    changed by mprotect()).  Two threads fault on each VMA concurrently, both
    calling __anon_vma_prepare() with only mmap_lock held for reading.
    
        THREAD A                             THREAD B
    __anon_vma_prepare                __anon_vma_prepare
     find_mergeable_anon_vma() -> NULL
     anon_vma = anon_vma_alloc();
       anon_vma->root = anon_vma;
     // the two stores may be reordered
     vma->anon_vma = anon_vma;
                                       // finds A's anon_vma
                                       anon_vma = find_mergeable_anon_vma(vma);
                                       anon_vma_lock_write(anon_vma);
                                         // may still see the old root
                                         down_write(&anon_vma->root->rwsem);
                                       anon_vma_unlock_write(anon_vma);
                                         // see the new root, never unlock old
                                         up_write(&anon_vma->root->rwsem);
    
    thread A triggers page fault and calls __anon_vma_prepare() to prepare
    anon_vma for the faulting vma.  __anon_vma_prepare() allocates and
    initializes a new anon_vma, and then publishes it to the vma with a plain
    store.  anon_vma_prepare() only requires the mmap_lock to be held for
    reading, so two threads can fault on adjacent VMAs at the same time.
    While thread A publishes a new anon_vma, thread B could find the anon_vma
    via find_mergeable_anon_vma() and then locks anon_vma->root->rwsem.
    
    The store to anon_vma->root in anon_vma_alloc() and the store to
    vma->anon_vma can be reordered.  The anon_vma_lock_write() and spin_lock()
    only provide acquire semantics, which do not prevent prior stores from
    being reordered after them.  The release semantics of the corresponding
    spin_unlock() and anon_vma_unlock_write() come too late, the store to
    vma->anon_vma is already published before they take effect.  As a result,
    thread B can observe the following order:
    
        vma->anon_vma = anon_vma;
        anon_vma->root = anon_vma;
    
    The anon_vma slab is SLAB_TYPESAFE_BY_RCU, so a newly allocated anon_vma
    may reuse memory from a previously freed one.  The constructor
    (anon_vma_ctor) does not reset anon_vma->root, and __put_anon_vma()
    doesn't clear it either, so the old root value persists until
    anon_vma_alloc() overwrites it.  If that store isn't visible, thread B
    reads a root that points to the old anon_vma and locks it.
    
    As a result, thread B can call anon_vma_lock_write() with the old root,
    and call anon_vma_unlock_write() with the new root, leading to an anon_vma
    will never be unlocked, and another anon_vma couldn't be locked anymore
    (its count is dropped from 0 to -1 due to wrong unlock).
    
    To fix it, change the plain store `vma->anon_vma = anon_vma` to store
    release, so that the fields of anon_vma are visible before anon_vma is
    published to vma->anon_vma.
    
    At read side, the load of anon_vma and anon_vma->root have address
    dependency.  According to Documentation/memory-barriers.txt and some
    investigations, only Alpha needs address-dependency barriers and it has
    been handled by READ_ONCE() in reusable_anon_vma().
    
    We reproduced this issue in v5.10 with KSM enabled.  The kernel doesn't
    merge commit cf7e7a3503df ("mm: prevent KSM from breaking VMA merging for
    new VMAs"), so there are many adjacent VMAs that aren't merged but are
    compatible for anon_vma.
    
    Without this fix, our production environment could reproduce this issue
    about 2-5 times each month.  After adding a smp_mb() before
    anon_vma_lock_write(anon_vma) in __anon_vma_prepare(), which is different
    to this patch, this issue hasn't been reproduced for one month.
    
    Link: https://lore.kernel.org/[email protected]
    Fixes: 5c341ee1dfc8 ("mm: track the root (oldest) anon_vma")
    Signed-off-by: Jinjiang Tu <[email protected]>
    Signed-off-by: Andrew Morton <[email protected]>
    Reviewed-by: Lance Yang <[email protected]>
    Reviewed-by: Lorenzo Stoakes (ARM) <[email protected]>
    Acked-by: David Hildenbrand (Arm) <[email protected]>
    Acked-by: Vlastimil Babka (SUSE) <[email protected]>
    Cc: Minchan Kim <[email protected]>
    Cc: Harry Yoo <[email protected]>
    Cc: Hiroyouki Kamezawa <[email protected]>
    Cc: Jann Horn <[email protected]>
    Cc: Jinjiang Tu <[email protected]>
    Cc: Kefeng Wang <[email protected]>
    Cc: Larry Woodman <[email protected]>
    Cc: Liam R. Howlett <[email protected]>
    Cc: Nanyong Sun <[email protected]>
    Cc: Rik van Riel <[email protected]>
    Cc: <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
mptcp: return sk_wait_data() errors from recvmsg() [+ + +]
Author: Mark Amirkan <[email protected]>
Date:   Sun Sep 13 10:30:05 2026 +0000

    mptcp: return sk_wait_data() errors from recvmsg()
    
    [ Upstream commit 60404266ef3e0a1cd8f7a164060e0c83efb72f4b ]
    
    Commit 581302298524 ("mptcp: error out earlier on disconnect") made
    mptcp_recvmsg() stop when sk_wait_data() returns an error.  The error is
    stored in err, but the function then jumps to a path which returns
    copied.  When no data was copied, recvmsg() therefore returns zero and
    reports a false EOF.
    
    Store the result in copied, which is the value returned by the function.
    This also keeps the usual partial-read result when data was copied before
    the error.
    
    A recvmsg() blocked in one thread reproduces the issue when another
    thread disconnects the same MPTCP socket with connect(AF_UNSPEC).
    Before this change recvmsg() returns zero; afterwards it returns -EPIPE.
    
    Fixes: 581302298524 ("mptcp: error out earlier on disconnect")
    Cc: [email protected]
    Signed-off-by: Mark Amirkan <[email protected]>
    Reviewed-by: Matthieu Baerts (NGI0) <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
net/mlx5: Bridge, don't fail switchdev events of sibling eswitch ports [+ + +]
Author: Bernardo Soares <[email protected]>
Date:   Fri Sep 18 10:59:30 2026 +0100

    net/mlx5: Bridge, don't fail switchdev events of sibling eswitch ports
    
    [ Upstream commit 35e6f970f553954d92ba20afa885139a8e7dd0d7 ]
    
    mlx5 registers the bridge offload switchdev notifiers once per eswitch
    instance, but the notifier chains are global, so every instance sees
    every event and must filter out the ones that aren't its own. The
    existing filter, mlx5_esw_bridge_dev_same_hw(), only checks that the
    event netdevice sits on the same HCA - intentional for merged eswitch,
    where one bridge can span representors of several eswitches on one
    HCA - but same-HCA doesn't mean the instance actually has that port:
    peer ports are only created reactively from NETDEV_CHANGEUPPER, so an
    instance brought up after a sibling PF's port was already enslaved has
    none. The port object and attribute handlers claim the event anyway
    once same-HW passes, then fail the port lookup and return -EINVAL,
    which gets reported to user space even though the owning instance
    already handled it (e.g. "bridge vlan add ... RTNETLINK answers:
    Invalid argument"). Fix by filtering on the tracked port instead.
    
    The same gap exists in the generic recursive lower-device walk used by
    attribute changes on a bridge with more than one representor enslaved
    directly: mlx5_esw_bridge_lower_rep_vport_num_vhca_id_get() is entered
    with the bridge master netdevice, falls through to its generic
    netdev_for_each_lower_dev() loop, and returns as soon as the recursion
    into any one lower device yields a non-NULL rep - the underlying base
    case, mlx5_esw_bridge_rep_vport_num_vhca_id_get(), only checks
    mlx5_esw_bridge_dev_same_hw(), not ownership by the calling instance's
    br_offloads. mlx5_esw_bridge_lag_rep_get(), used for the LAG-master
    case, already filters on mlx5_esw_bridge_dev_same_esw() per candidate
    and so cannot select a sibling's rep; it is not the source of this bug.
    On a merged-eswitch HCA with a bridge spanning representors of more
    than one eswitch instance directly, the walk can return a sibling's rep
    instead of continuing to the one the calling instance actually owns, so
    the attribute change fails the same way as above. Fix by checking
    mlx5_esw_bridge_port_exists() at the point each rep is picked, same as
    the previous fix did for the notifier filter.
    
    Fixes: c358ea1741bc ("net/mlx5: Bridge, allow merged eswitch connectivity")
    Signed-off-by: Bernardo Soares <[email protected]>
    Cc: Vlad Buslov <[email protected]>
    Cc: Saeed Mahameed <[email protected]>
    Reviewed-by: Mark Bloch <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net/mlx5: Bridge, don't fail unlink of untracked/unsupported peer ports [+ + +]
Author: Bernardo Soares <[email protected]>
Date:   Fri Sep 18 10:59:31 2026 +0100

    net/mlx5: Bridge, don't fail unlink of untracked/unsupported peer ports
    
    [ Upstream commit 2e51097c982b7b22382fda3202ba29f0ea33e8c0 ]
    
    mlx5_esw_bridge_vport_unlink() returns -EINVAL when the port isn't
    tracked by this instance's br_offloads. This is reachable on a sibling
    instance that registered its notifier after the port was already
    enslaved: it never saw the NETDEV_CHANGEUPPER link event, so
    peer_link() never created a peer port for it, but it does see the
    later unlink event and fails. Return 0 instead, and give
    mlx5_esw_bridge_vport_peer_unlink() the same merged_eswitch capability
    guard peer_link() already has, since without it peer_link() likewise
    never creates a port to unlink.
    
    This also matters beyond the -EINVAL itself:
    mlx5_esw_bridge_switchdev_port_event() runs on the per-netns
    netdev_chain, and notifier_from_errno(-EINVAL) sets NOTIFY_STOP_MASK,
    which call_netdevice_notifiers_info() checks to stop calling further
    listeners on that chain - so the old -EINVAL silently dropped the
    event for any listener registered later on the same chain, even
    though none of it was visible to user space since
    __netdev_upper_dev_unlink() discards the return value.
    
    Fixes: c358ea1741bc ("net/mlx5: Bridge, allow merged eswitch connectivity")
    Signed-off-by: Bernardo Soares <[email protected]>
    Reviewed-by: Mark Bloch <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net/mlx5: devcom, Base component size on linked devices [+ + +]
Author: Shay Drory <[email protected]>
Date:   Tue Sep 15 14:34:57 2026 +0300

    net/mlx5: devcom, Base component size on linked devices
    
    [ Upstream commit d09e8f64653c93da5793c16be19330968f2a32e6 ]
    
    mlx5_devcom_comp_get_size() returns the component's kref count. That
    kref is bumped in mlx5_devcom_register_component() under comp_list_lock,
    before the comp_dev is linked onto comp_dev_list_head under comp->sem.
    The event broadcast (mlx5_devcom_locked_send_event()) walks that list.
    
    Hence, a caller can read the expected size, but send_event won't be sent
    to all peers. In the SD group registration path, this lets a member
    broadcast its role-election event over an incomplete list, electing a
    primary that never completes the group, is never marked ready, and
    leaves the group with a stale primary.
    
    Track the number of linked comp_devs in a dedicated counter, maintained
    under comp->sem together with the list add/remove, and return it from
    mlx5_devcom_comp_get_size().
    
    Fixes: 9bb1ac80738a ("net/mlx5: devcom, Add component size getter")
    Signed-off-by: Shay Drory <[email protected]>
    Reviewed-by: Akiva Goldberger <[email protected]>
    Signed-off-by: Tariq Toukan <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net/mlx5: Fix rev_entry reference leak in mlx5_tc_ct_shared_counter_get() [+ + +]
Author: Wentao Liang <[email protected]>
Date:   Thu Sep 17 11:31:31 2026 +0000

    net/mlx5: Fix rev_entry reference leak in mlx5_tc_ct_shared_counter_get()
    
    commit 0bf6bb567f0edaa771e7dd208ef98da50e6a4485 upstream.
    
    When the reverse entry is found but its counter is already being
    released, refcount_inc_not_zero() fails and the reference taken by
    mlx5_tc_ct_entry_get() is never dropped before falling through to
    create_counter.  Drop it so the reverse entry is not kept alive forever
    by a shared counter lookup that did not use it.
    
    Fixes: 1edae2335adf ("net/mlx5e: CT: Use the same counter for both directions")
    Cc: [email protected]
    Signed-off-by: Wentao Liang <[email protected]>
    Reviewed-by: Tariq Toukan <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net/mlx5: LAG, reload IB reps of LAG master before the rest [+ + +]
Author: Shay Drory <[email protected]>
Date:   Tue Sep 15 14:34:59 2026 +0300

    net/mlx5: LAG, reload IB reps of LAG master before the rest
    
    [ Upstream commit bae23d1ae62092c7f0ec6d5f7e1be5d164822638 ]
    
    In a shared-FDB LAG the master device creates the bond IB device; the
    other LAG members do not create their own, they populate a port inside
    the master's IB device. mlx5_lag_reload_ib_reps_unlocked() reloaded the
    members' IB reps in iteration order, with no guarantee the master is
    reloaded first. When a non-master member is reloaded before the master,
    it tries to populate its port in an IB device that has not been
    recreated yet.
    
    Hence, reload the master's IB reps first, then every other member.
    
    Fixes: 2b204cdb1206 ("net/mlx5: LAG, use xa_alloc to manage LAG device indices")
    Signed-off-by: Shay Drory <[email protected]>
    Reviewed-by: Akiva Goldberger <[email protected]>
    Signed-off-by: Tariq Toukan <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net/mlx5: SD, unload reps on shared FDB create error path [+ + +]
Author: Shay Drory <[email protected]>
Date:   Tue Sep 15 14:34:58 2026 +0300

    net/mlx5: SD, unload reps on shared FDB create error path
    
    [ Upstream commit e1e29ada2b938b13ba689a06a8bd8604564da2b3 ]
    
    mlx5_lag_shared_fdb_create() sets sd_fdb_active on every group member
    before reloading the representors, so mlx5_lag_is_active() is already
    true and the guard in mlx5_esw_offloads_rep_load() does not skip the
    VF/SF reps. If the reload then fails, the error path clears
    sd_fdb_active and destroys the shared FDB, leaving the reps loaded
    while SD LAG is inactive - the state cited commit was written
    to prevent.
    
    Unload the reps in the error path as well.
    
    Fixes: 68c2dd59a6c7 ("net/mlx5: E-Switch, Tie rep load/unload to SD LAG state")
    Signed-off-by: Shay Drory <[email protected]>
    Reviewed-by: Akiva Goldberger <[email protected]>
    Signed-off-by: Tariq Toukan <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
net/mlx5e: advertise MACsec offload only when supported [+ + +]
Author: Ralf Lici <[email protected]>
Date:   Thu Sep 17 14:27:23 2026 +0200

    net/mlx5e: advertise MACsec offload only when supported
    
    commit 4581c3d2adc3c73a019bc38db64ca11f28bbd7fd upstream.
    
    Commit 339ccec8d43d ("net/mlx5: Enable MACsec offload feature for VLAN
    interface") added NETIF_F_HW_MACSEC unconditionally to vlan_features so
    that VLAN devices could inherit MACsec offload support.
    
    mlx5e_build_nic_netdev subsequently copies vlan_features into
    hw_features and features. As a result, all mlx5e NIC netdevices
    advertise MACsec hardware offload, even when the firmware does not
    support it and the driver does not install macsec_ops.
    
    Set the MACsec feature bits in mlx5e_macsec_build_netdev, after device
    capabilities have been validated. This preserves MACsec-over-VLAN
    support and the ethtool feature control on capable devices, without
    advertising either on unsupported hardware.
    
    Fixes: 339ccec8d43d ("net/mlx5: Enable MACsec offload feature for VLAN interface")
    Cc: [email protected]
    Reviewed-by: Tariq Toukan <[email protected]>
    Signed-off-by: Ralf Lici <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net/mlx5e: fix swapped IPv6 IPsec policy masks [+ + +]
Author: Andrea Parri <[email protected]>
Date:   Thu Sep 17 13:55:42 2026 +0200

    net/mlx5e: fix swapped IPv6 IPsec policy masks
    
    commit 10de7ed8ef4840da9ca21de4c29578657ac367db upstream.
    
    IPv6 XFRM policies may use different source and destination prefix
    lengths. mlx5e_ipsec_policy_mask() builds the corresponding masks
    independently, but setup_fte_addr6() installs each mask in the opposite
    address field.
    
    When the prefix lengths differ, this makes the source match use the
    destination prefix and the destination match use the source prefix. The
    resulting hardware rule can both miss traffic covered by the policy and
    match traffic outside it.
    
    Install each mask in its corresponding match field.
    
    Fixes: ca7992f52c2c ("net/mlx5e: Properly match IPsec subnet addresses")
    Cc: [email protected]
    Signed-off-by: Andrea Parri <[email protected]>
    Reviewed-by: Tariq Toukan <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
net/rds: size a connection's path set by the transport it ends up with [+ + +]
Author: Allison Henderson <[email protected]>
Date:   Mon Sep 21 14:50:27 2026 -0700

    net/rds: size a connection's path set by the transport it ends up with
    
    [ Upstream commit ba3d1f480c7a3fba963e7867ad6cc557c197acbc ]
    
    __rds_conn_create() computes npaths from the caller's transport before
    it decides whether a connection to one of the host's own addresses is
    to be handled by the loopback transport instead.  That substitution is
    what an RDS/TCP socket sending to a local address gets, and after it
    the path init loop still runs for the TCP transport's RDS_MPATH_WORKERS
    paths and allocates an ordered workqueue for each, while
    rds_loop_conn_alloc() only ever provides transport data for path 0.
    
    rds_conn_destroy() sizes its teardown from c_trans, by then the
    loopback transport, so it visits path 0 only - and
    rds_conn_path_destroy() would skip the other paths anyway, since it
    returns before destroy_workqueue() for a path without transport data.
    kfree(c_path) then drops the last pointers to seven workqueues.  That
    repeats for every such connection, on every netns teardown or module
    unload, and every distinct local destination address is a separate
    connection.
    
    Recompute npaths once the transport is final, so that creation and
    destruction agree on the set of paths.  The c_path array stays sized
    for the caller's transport; the unused entries are freed with it.
    
    Fixes: 4716af3897e9 ("net/rds: Give each connection path its own workqueue")
    Signed-off-by: Allison Henderson <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
net/sched: act_ct: avoid modifying shared unconfirmed ct entry [+ + +]
Author: Ilya Maximets <[email protected]>
Date:   Tue Sep 29 20:04:35 2026 -0400

    net/sched: act_ct: avoid modifying shared unconfirmed ct entry
    
    [ Upstream commit f85009dfcd65e5969526b0db7a49b5413746e630 ]
    
    In a case where skb with an unconfirmed ct entry gets cloned, we may
    end up processing both again but with different sets of extensions.
    
    The series of events:
    
     1. The first clone wants to commit and runs the helpers wiring up
        the extension pointer into the expectation list.
     2. Then it looses the confirmation keeping the entry unconfirmed.
     3. Second clone now wants to commit labels or run NAT and adds the
        new extension for that breaking the pointer in the expectation
        list causing UAF on the destruction path later.
    
    While this is possible to trigger, there should be no practical
    network pipeline where we need to process both clones without
    modifications in the same zone.  So, let's just reset the entry in
    case for some reason we got an skb with a shared one.  This doesn't
    affect any known use cases, but avoids any potential problems with
    sharing and modification of the unconfirmed ct entry.
    
    Unlike openvswitch module, act_ct allows for NAT without commit.
    Changing that would be a uAPI break.  So, act_ct needs to reset on NAT
    regardless of the commit flag to avoid reallocation of the extension
    space.  This, however, doesn't really change the picture for sensible
    networking cases as there should be no need to run the same packet
    twice (before and after the clone) through conntrack without packet
    header or zone changes and without commit.
    
    The fixes tag points to the introduction of helpers, since that's the
    main UAF trigger for the sharing.
    
    Fixes: a21b06e73191 ("net: sched: add helper support in act_ct")
    Cc: [email protected]
    Reported-by: Axel Mierczuk <[email protected]>
    Signed-off-by: Ilya Maximets <[email protected]>
    Reviewed-by: Aaron Conole <[email protected]>
    Reviewed-by: Xin Long <[email protected]>
    Reviewed-by: Jamal Hadi Salim <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net/sched: act_ct: don't WARN on benign flow_offload_alloc() failure [+ + +]
Author: Nguyen Ngoc Thang <[email protected]>
Date:   Tue Sep 15 22:08:16 2026 +0700

    net/sched: act_ct: don't WARN on benign flow_offload_alloc() failure
    
    [ Upstream commit 47abe7a5c4eb53269aca3506446f851572a059a3 ]
    
    flow_offload_alloc() returns NULL when the conntrack entry is dying
    (e.g. raced with a conntrack flush) or when the GFP_ATOMIC allocation
    fails; both are expected under load and neither is a kernel bug. This
    path runs from softirq on every committed packet, so with
    panic_on_warn=1 an unprivileged user can panic the box just by racing
    a conntrack flush against a `tc ... action ct commit` classifier.
    
    Reproduced with a custom repro under QEMU: a small, fixed set of UDP
    flows through `tc filter ... action ct commit` on lo, raced against
    threads flooding bare ctnetlink CT_DELETE (flush) requests. Hits
    WARNING: net/sched/act_ct.c:437 (tcf_ct_flow_table_add(), inlined
    into tcf_ct_act() in this build) within ~15s on the unpatched kernel;
    same setup is clean on the patched kernel. The fix itself is
    behavior-preserving: both branches already did `goto err_alloc`
    before and after, only the WARN is removed.
    
    Fixes: 64ff70b80fd4 ("net/sched: act_ct: Offload established connections to flow table")
    Reported-by: [email protected]
    Closes: https://syzkaller.appspot.com/bug?extid=6cc37aba98dac721c415
    Signed-off-by: Nguyen Ngoc Thang <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net/sched: act_ct: fix helper UAF due to extensions realloc [+ + +]
Author: Ilya Maximets <[email protected]>
Date:   Tue Sep 29 20:04:14 2026 -0400

    net/sched: act_ct: fix helper UAF due to extensions realloc
    
    [ Upstream commit dad19b59da050cb60d3f7023dac2a042a84bf0bd ]
    
    While calling the helpers, a raw pointer to the extensions area is
    wired into expectations list:
    
      -> nf_ct_helper()
       -> helper->help()
        -> nf_ct_expect_related_report()
         -> nf_ct_expect_insert()
          -> hlist_add_head_rcu(&exp->lnode, &master_help->expectations)
    
    In case the connection is not confirmed yet, more extensions can be
    added afterwards with *_ext_add() calls reallocating the extension
    space and leaving the now invalid pointer in the expectations list
    that is later accessed while removing the expectation.
    
    Make sure that helpers are called at the end after all the other
    extensions are already added.
    
    Note that the helper rejection now leaves the mark and labels set,
    but that's not different from how the NAT was handled before or how
    the mark and the labels were handled on confirmation failure.  And
    there are no atomicity guarantees provided by the API anyway.
    
    Fixes: a21b06e73191 ("net: sched: add helper support in act_ct")
    Cc: [email protected]
    Reported-by: Axel Mierczuk <[email protected]>
    Signed-off-by: Ilya Maximets <[email protected]>
    Reviewed-by: Xin Long <[email protected]>
    Reviewed-by: Jamal Hadi Salim <[email protected]>
    Reviewed-by: Aaron Conole <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    [ Preserved the existing add_helper condition that upstream had removed. ]
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net/sched: act_gate: budget the per-entry list in get_fill_size [+ + +]
Author: Victor Nogueira <[email protected]>
Date:   Sun Sep 20 14:07:01 2026 -0300

    net/sched: act_gate: budget the per-entry list in get_fill_size
    
    [ Upstream commit cfa165cbfbed9d0f4bbc22fef4309f595a3ab187 ]
    
    tcf_gate_get_fill_size returns only the TCA_GATE_PARMS size, but
    tcf_gate_dump also emits three 64-bit timestamps, the clock id, flags,
    priority and the variable-length TCA_GATE_ENTRY_LIST nest. The per-entry
    nest is unbounded: parse_gate_list places no cap on the number of
    sched-entries, so a gate with many entries can push the real dump well
    past the skb that tca_get_fill allocates from this size.
    
    RTM_NEWACTION then fails the add-notify with -EINVAL while the action is
    already committed to the IDR, and a subsequent RTM_GETACTION on the
    installed gate also returns -EINVAL because its dump no longer fits.
    
    Fix this by accounting for the missing fields in tcf_gate_get_fill_size
    along with all elements in the entries list.
    
    Note that sizing the reply from the action lets an oversized gate
    install cleanly for the first time: with the input unbounded by
    parse_gate_list, the sized skb can now grow well above
    NLMSG_GOODSIZE per netlink request (a transient GFP_KERNEL allocation
    reachable only with namespace-local CAP_NET_ADMIN). Overload from a
    malicious netns admin is hardening material, not net, per the
    discussion at
    https://lore.kernel.org/netdev/[email protected]/;
    a follow-up patch for net-next will cap the sched-entry count.
    
    Fixes: 4e76e75d6aba ("net sched actions: calculate add/delete event message size")
    Reported-by: Sashiko <[email protected]>
    Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/[email protected]
    Tested-by: hybris <[email protected]>
    Co-developed-by: Jamal Hadi Salim <[email protected]>
    Signed-off-by: Jamal Hadi Salim <[email protected]>
    Signed-off-by: Victor Nogueira <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net/sched: act_ife: validate metadata length before decoding [+ + +]
Author: Fang Xieyan <[email protected]>
Date:   Mon Sep 21 20:54:41 2026 +0800

    net/sched: act_ife: validate metadata length before decoding
    
    [ Upstream commit d6ec384c87cc851cfd13bb18c99ce351ccee6192 ]
    
    skbmark_decode(), skbprio_decode() and skbtcindex_decode() read fixed-size
    values from the TLV payload without validating its length.
    
    A malformed IFE frame can declare a shorter payload, causing the decoders
    to consume bytes beyond the declared metadata value:
    
    [TLV type=IFE_META_SKBMARK len=4]
    -> dlen == 0, but decode reads 4 bytes
    
    The decoder may therefore set skb metadata from unintended input.
    
    Validate the payload length before decoding and return -EINVAL for
    invalid lengths. Read the values with get_unaligned_be32() and
    get_unaligned_be16(), as TLV payloads are not guaranteed to be
    aligned. Teach tcf_ife_decode() to log a decoder error separately
    from an unknown metaid; both are counted as overlimits and decoding
    continues with the remaining metadata.
    
    The metadata length issue was found by an automated audit of the IFE
    decode path at v6.18-rc7 and reproduced with a userspace sanitizer
    model of the decode path. Compile-tested on x86_64 with defconfig and
    NET_ACT_IFE=y: act_ife.o and the three act_meta_*.o build
    warning-free.
    
    Fixes: 084e2f6566d2 ("Support to encoding decoding skb mark on IFE action")
    Fixes: 200e10f46936 ("Support to encoding decoding skb prio on IFE action")
    Fixes: 408fbc22ef1e ("net sched ife action: Introduce skb tcindex metadata encap decap")
    Assisted-by: Hawkeye:GLM-5.3-flash
    Assisted-by: Qoder:Qwen3.8-Max
    Signed-off-by: Fang Xieyan <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net/sched: cls_u32: fix manual hash table handle IDR aliasing [+ + +]
Author: Jamal Hadi Salim <[email protected]>
Date:   Wed Sep 16 06:01:14 2026 -0400

    net/sched: cls_u32: fix manual hash table handle IDR aliasing
    
    [ Upstream commit 0a5f5d9e94dead312d32c366b917c64e552b72f7 ]
    
    A u32 hash table created with an explicit handle ('tc filter add ...
    handle 801: u32 divisor N') keys its IDR entry on the raw handle, while
    the destroy paths free it under handle2id(handle). The two key domains
    disagree for handles in the 0x800..0xFFF htid range:
    handle2id() folds them back into the auto-allocated id space (1..0x7FF).
    
    A manual table therefore leaves its raw-keyed IDR entry unreachable on
    delete (a permanent leak), and its delete can drop the idr entry of an
    unrelated live auto table. A later auto allocation can then hand out a
    handle that aliases the live manual table; u32_lookup_ht() first-match
    routes lookups and TCA_U32_LINK for that htid to the wrong table.
    
    Key the divisor-path alloc on handle2id(handle) so allocation and
    removal share one key domain. A manual handle that maps onto an id
    already in use is rejected with -ENOSPC, and auto allocation skips ids
    held by live manual tables.
    
    Conditions to recreate:
      ip link add test0 type dummy
      tc qdisc add dev test0 clsact
      tc filter add dev test0 ingress protocol ip pref 1 \
              handle 801: u32 divisor 16
      tc filter add dev test0 ingress protocol ip pref 2 u32 divisor 16
      tc -d filter show dev test0 ingress | grep 'fh 801:'
      # unpatched: two live tables with handle 0x80100000 (the pref 2 root
      # hnode is auto-allocated id 1); patched: the auto hnode takes id 2.
    
    Also tested with a poc with a live u32 table on the block, add/delete a manual
    table 'handle 901: u32 divisor 1' twice; unpatched, the re-add fails with
    -ENOSPC because the raw key leaked on the first delete.
    
    Fixes: 73af53d82076 ("net: sched: cls_u32: Fix u32's systematic failure to free IDR entries for hnodes.")
    Reported-by: Sashiko (gemini + nipa) <[email protected]>
    Closes: https://sashiko.dev/#/patchset/[email protected]
    Reviewed-by: Victor Nogueira <[email protected]>
    Tested-by: hybris <[email protected]>
    Signed-off-by: Jamal Hadi Salim <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net/sched: fix potential stack infoleak in em_text_dump() [+ + +]
Author: Bernard Ladenthin <[email protected]>
Date:   Fri Sep 18 15:39:37 2026 +0200

    net/sched: fix potential stack infoleak in em_text_dump()
    
    [ Upstream commit 9c572a83037a7dcd653ba3a9cc468c16b857d0c9 ]
    
    em_text_dump() allocates struct tcf_em_text on the stack without zeroing
    it.  strscpy() writes the algorithm name and a NUL terminator into
    conf.algo[], leaving the remaining bytes uninitialised.  nla_put_nohdr()
    then copies the full struct to the netlink response.
    
    KMSAN on Linux 7.2-rc6 reports two kernel-infoleak splats from this path,
    one triggered via "tc filter show" and one via a raw RTM_GETTFILTER dump:
    
      BUG: KMSAN: kernel-infoleak in _copy_to_iter+0x1c9/0x2620
        nla_put_nohdr+0x83/0x130
        em_text_dump+0x291/0x550
      Local variable conf created at: em_text_dump+0x5d/0x550
      Bytes 168-179 of 199 are uninitialized
    
    I am not certain whether this constitutes a real security problem in
    practice: the test was conducted in a controlled KMSAN environment and
    the leaked stack bytes may or may not carry sensitive data on actual
    production kernels.  I am reporting it because KMSAN flagged it as a
    kernel-infoleak and the fix is straightforward.  I can provide a
    userspace reproducer on request.
    
    The original code used strncpy() which zero-pads to the destination size.
    Commit b04202d6065c ("net/sched: replace strncpy with strscpy") replaced
    it with strscpy(), which does not pad, creating this condition.
    Zero-initialising the struct closes it.
    
    Fixes: b04202d6065c ("net/sched: replace strncpy with strscpy")
    Link: https://lore.kernel.org/netdev/[email protected]/
    Assisted-by: Claude:claude-sonnet-4-6 [KMSAN]
    Signed-off-by: Bernard Ladenthin <[email protected]>
    Acked-by: Jamal Hadi Salim <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net/sched: reject IDR error pointers when deleting actions [+ + +]
Author: Weiming Shi <[email protected]>
Date:   Mon Sep 14 14:51:23 2026 +0800

    net/sched: reject IDR error pointers when deleting actions
    
    commit c82b797abe668d0b668601a93ba2c0b071a63574 upstream.
    
    tcf_action_delete() drops the reference held by its lookup before calling
    tcf_idr_delete_index() with the saved action index.  An unlocked
    classifier can remove that action and reserve the same IDR slot with
    ERR_PTR(-EBUSY) in between.
    
    tcf_idr_delete_index() only checks the lookup result for NULL.  It
    therefore treats the reservation as a tc_action and dereferences
    tcfa_bindcnt.  A hardware execution breakpoint was used to schedule the
    interleaving without changing the kernel source.  KASAN reported this
    decoded trace:
    
      BUG: KASAN: null-ptr-deref in tca_action_gd+0x5b9/0x1010
      Read of size 4 at addr 0000000000000010 by task poc/150
      Oops: general protection fault, probably for non-canonical address 0xdffffc0000000002
      RIP: tca_action_gd+0x5c0/0x1010:
        arch_atomic_read at arch/x86/include/asm/atomic.h:23
        raw_atomic_read at include/linux/atomic/atomic-arch-fallback.h:457
        atomic_read at include/linux/atomic/atomic-instrumented.h:33
        tcf_idr_delete_index at net/sched/act_api.c:766
        tcf_action_delete at net/sched/act_api.c:1859
        tcf_del_notify at net/sched/act_api.c:2014
        tca_action_gd at net/sched/act_api.c:2064
      R13: 0000000000000010 R15: fffffffffffffff0
      Kernel panic - not syncing: Fatal exception
    
    R15 contains ERR_PTR(-EBUSY), and adding the tcfa_bindcnt offset produces
    the address in R13.  With the guard applied, the same reproducer returned
    -ENOENT without a KASAN report or panic.  Treat error pointers as absent
    and return -ENOENT.
    
    Fixes: 0190c1d452a9 ("net: sched: atomically check-allocate action")
    Cc: [email protected]
    Reported-by: Xiang Mei <[email protected]>
    Signed-off-by: Weiming Shi <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net/sched: sch_hfsc: bound the classify inner-filter walk with a drift budget [+ + +]
Author: Jamal Hadi Salim <[email protected]>
Date:   Thu Sep 17 06:57:32 2026 -0400

    net/sched: sch_hfsc: bound the classify inner-filter walk with a drift budget
    
    [ Upstream commit 8a60ade2277e1f0e0d0578d565354e52292fa46d ]
    
    hfsc_classify() applies the "filter may only point downwards" level check
    only when the filter result carries no bound class. A filter created with
    a flowid gets res.class set once at bind time, so the check never runs for
    it during classification. hfsc_adjust_levels() can later raise a class's
    level without revalidating existing bindings, leaving two binds that were
    each legal at bind time pointing at each other; the classify walk then
    bounces between two interior classes forever with the qdisc lock held and
    BH disabled — a soft lockup from a single packet. The stuck walk trips
    the watchdog:
    
      watchdog: BUG: soft lockup - CPU#3 stuck for 13s! [ping:444]
      RIP: 0010:u32_classify+0x542/0x17f0
      ...
      tcf_classify+0x66/0xa0
      hfsc_enqueue+0x166/0xdf0
    
    Bound the traversal with a budget of non-descending hops, the only way a
    configured walk can move without descending the class tree once levels
    drift after bind time. The budget is cumulative over the whole walk and
    is deliberately not reset on a descending hop: a chain that alternates a
    descent with a lateral hop would return the budget every lap and never
    trip. Descending hops never decrement it, so legitimately deep trees are
    unaffected and a terminating lateral chain still classifies normally.
    Drop the packet with a rate-limited warning once the budget is exhausted,
    mirroring the merged HTB fix.
    
    This is a follow-up to commit 729c4896ab82 ("net/sched: sch_htb: limit
    htb_classify inner-class filter hops"), which bounded the same classify
    loop on the HTB side but left the HFSC walk unbounded.
    
    Conditions to recreate the bug:
    - CONFIG_NET_SCHED, CONFIG_NET_SCH_HFSC, CONFIG_NET_CLS_U32,
      CONFIG_LOCKUP_DETECTOR.
    - Build a cycle with two legal-at-bind-time flowid binds and a level
      drift: class X 1:1 (child of root) with leaf child 1:10; class Y 1:2
      (sibling of X) with children 1:20 and 1:200; root u32 filter flowid
      1:1; filter on X flowid 1:2 (legal when Y is a leaf); after Y's level
      rises to 2, filter on Y flowid 1:1 (legal then). Send one packet (ping
      on the device). Unfixed kernel: classify spins with the qdisc lock
      held; with softlockup_panic=1 it panics.
    - Reachable from unprivileged user via unshare -Urn (CAP_NET_ADMIN).
    
    Fixes: a2f79227138c ("net_sched: sch_hfsc: fix classification loops")
    Reported-by: Sashiko (gemini + nipa) <[email protected]>
    Closes: https://lore.kernel.org/netdev/[email protected]/
    Link: https://sashiko.dev/#/patchset/[email protected]
    Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/QDISC-CTUU.v2.20260913192614%40mojatatu.com
    Reviewed-by: Victor Nogueira <[email protected]>
    Tested-by: hybris <[email protected]>
    Signed-off-by: Jamal Hadi Salim <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net/sched: sch_teql: fix shadowed err in __teql_resolve() [+ + +]
Author: Eric Dumazet <[email protected]>
Date:   Thu Sep 24 08:29:50 2026 +0000

    net/sched: sch_teql: fix shadowed err in __teql_resolve()
    
    [ Upstream commit 907b978e82cb4c1c245fc2985bb27c5d5c88c8f6 ]
    
    __teql_resolve() declares an inner 'int err;' inside the
    'if (neigh_event_send(n, skb_res) == 0)' block, shadowing the outer
    'int err = 0;'. As a result, a negative return from dev_hard_header()
    is written to the inner variable and __teql_resolve() still returns 0.
    
    Remove the shadowed variable and set the outer err to -EINVAL when
    dev_hard_header() returns a negative error.
    
    Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
    Closes: https://lore.kernel.org/netdev/[email protected]/
    Cc: Jamal Hadi Salim <[email protected]>
    Cc: Jiri Pirko <[email protected]>
    Signed-off-by: Eric Dumazet <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
net/smc: fix UAF on lgr list traversal in smcr_port_err() [+ + +]
Author: Sidraya Jayagond <[email protected]>
Date:   Tue Sep 22 09:31:49 2026 +0200

    net/smc: fix UAF on lgr list traversal in smcr_port_err()
    
    [ Upstream commit 61cb282fe97b3b0ba32ca09417a693162bf4ae3f ]
    
    smcr_port_err() traverses smc_lgr_list.list without holding
    smc_lgr_list.lock, allowing a concurrent smc_lgr_terminate_sched()
    to free an lgr while it is still being dereferenced.
    
    Hold smc_lgr_list.lock across the traversal. Update
    smc_ib_gid_check() to call smcr_port_err() after releasing the lock.
    
    Fixes: 541afa10c126 ("net/smc: add smcr_port_err() and smcr_link_down() processing")
    Reviewed-by: Mahanta Jambigi <[email protected]>
    Signed-off-by: Sidraya Jayagond <[email protected]>
    Reviewed-by: Dust Li <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
net: airoha: npu: cancel wdt_work after releasing the WDT IRQ [+ + +]
Author: Myeonghun Pak <[email protected]>
Date:   Mon Sep 21 20:09:14 2026 -0400

    net: airoha: npu: cancel wdt_work after releasing the WDT IRQ
    
    commit 4bdee8060d1e4581624e68fbd369b1afb14df4bc upstream.
    
    airoha_npu_remove() calls cancel_work_sync() on each core's wdt_work,
    but the watchdog IRQ that queues it is requested with devm_request_irq()
    and is freed only after .remove() returns.  airoha_npu_wdt_handler() can
    therefore schedule_work() again once the cancel has returned.  struct
    airoha_npu, which contains the work, is devm_kzalloc()'d and is freed in
    that same unwind, so the late work dereferences freed memory.
    
    Register the work with devm_work_autocancel() before devm_request_irq()
    and drop .remove().  Devres runs in reverse order, so the IRQ is freed
    before cancel_work_sync(), including when probe fails.  A cancel left in
    .remove() cannot get that order.  Initializing the work first also stops
    a pending watchdog interrupt from queuing an uninitialized work item.
    Probe currently calls INIT_WORK() only after devm_request_irq().
    
    This issue was identified during our ongoing static-analysis research
    while reviewing kernel code.
    
    Fixes: 23290c7bc190 ("net: airoha: Introduce Airoha NPU support")
    Cc: [email protected] # 6.15+
    Co-developed-by: Ijae Kim <[email protected]>
    Signed-off-by: Ijae Kim <[email protected]>
    Signed-off-by: Myeonghun Pak <[email protected]>
    Acked-by: Lorenzo Bianconi <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: allow IFLA_INET_CONF messages when NLA_F_NESTED unset [+ + +]
Author: Quentin Armitage <[email protected]>
Date:   Tue Sep 15 22:33:21 2026 +0100

    net: allow IFLA_INET_CONF messages when NLA_F_NESTED unset
    
    [ Upstream commit 6c096bb08de97cdca051fecddad22cac6a1fd275 ]
    
    Commit fa8fca88714c ("ipv4: validate IPV4_DEVCONF attributes properly")
    added validation of IFLA_INET_CONF attributes, and in the process
    changed the call of nla_for_each_nested() to nla_parse_nested(). A
    side effect of this change is that the IFLA_INET_CONF option is now
    tested for NLA_F_NESTED being set, and fails if it is not. Prior to the
    commit there was no check of NLA_F_NESTED.
    
    Change nla_parse_nested() to nla_parse(). This restores the previous
    functionality of not checking NLA_F_NESTED, thereby allowing code that
    (incorrectly) doesn't set NLA_F_NESTED to continue to work.
    
    This issue was identified because keepalived started logging errors when
    it was configuring macvlans that it created.
    
    Fixes: fa8fca88714c ("ipv4: validate IPV4_DEVCONF attributes properly")
    Signed-off-by: Quentin Armitage <[email protected]>
    Reviewed-by: Ido Schimmel <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: arp: terminate device name before lookup [+ + +]
Author: Zijie Huang <[email protected]>
Date:   Mon Sep 21 01:36:19 2026 +0800

    net: arp: terminate device name before lookup
    
    commit d8b6529e80bcb4fb8177121404cbb3377acaebd2 upstream.
    
    The ARP ioctl copies a user-provided struct arpreq into a stack object. Its
    arp_dev field may contain IFNAMSIZ bytes without a NUL terminator.
    
    Such input is passed to dev_get_by_name_rcu() or __dev_get_by_name(), where
    strcmp() can read past the end of the stack object when a matching
    alternative interface name exists.
    
    Terminate the field before the lookup to prevent the out-of-bounds read.
    
    Fixes: 36fbf1e52bd3 ("net: rtnetlink: add linkprop commands to add and delete alternative ifnames")
    Cc: [email protected]
    Reported-by: Vega <[email protected]>
    Signed-off-by: Zijie Huang <[email protected]>
    Signed-off-by: Ren Wei <[email protected]>
    Reviewed-by: Ido Schimmel <[email protected]>
    Link: https://patch.msgid.link/fabf02a70787d17299e4b3153eadffaf20d154b3.1789910973.git.milkory@outlook.com
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: atl1: fix soft lockup on out-of-range cmb_tpd_next_to_clean read [+ + +]
Author: Gajdos Tamás <[email protected]>
Date:   Mon Sep 21 11:13:34 2026 +0200

    net: atl1: fix soft lockup on out-of-range cmb_tpd_next_to_clean read
    
    commit 43e746821f5f5afbbf68e388bf9fbe221e03bfca upstream.
    
    Same issue as atl1c (see the first commit in this series, "net:
    atl1c: fix soft lockup on out-of-range tpd_cons read"): the hardware
    can report an out-of-range cmb_tpd_next_to_clean (seen as 0xffff)
    while the PCIe link/MAC is resetting. An out-of-range value can
    never be reached and the loop below would spin forever. Treat it as
    "nothing new to clean" instead.
    
    Fixes: f3cc28c797604f ("Add Attansic L1 ethernet driver.")
    Cc: [email protected]
    Signed-off-by: Gajdos Tamás <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: atl1c: fix soft lockup on out-of-range tpd_cons read [+ + +]
Author: Gajdos Tamás <[email protected]>
Date:   Mon Sep 21 11:13:32 2026 +0200

    net: atl1c: fix soft lockup on out-of-range tpd_cons read
    
    commit 36c2009d90f2210ef92e6f4f2850e8b57b09e754 upstream.
    
    The hardware can report an out-of-range tpd_cons (seen as 0xffff)
    while the PCIe link/MAC is resetting. An out-of-range value can
    never be reached and the loop below would spin forever. To avoid
    a soft lockup treat it as "nothing new to clean" instead.
    
    Reproduced on two machines, same NIC (Qualcomm Atheros AR8151 v2.0,
    4-port), triggered by rebooting a Mikrotik CCR2004 PCIe card that
    the ports are directly linked to:
    
    - Ubuntu 26.04.1 LTS, kernel 7.0.0-31-generic. The link-flap
      precursor, before the lockup was captured with a full trace
      elsewhere:
    
        atl1c 0000:05:00.0 enp5s0f0: NETDEV WATCHDOG: CPU: 4: transmit queue 2 timed out 489984 ms
        atl1c 0000:05:00.0: MAC state machine can't be idle since disabled for 10ms second
        atl1c 0000:05:00.0: atl1c: enp5s0f0 NIC Link is Up<65535 Mbps Full Duplex>
    
      65535 (0xffff) here is the same value tpd_cons reads back once the
      loop below gets stuck.
    
    - Proxmox VE, kernel 7.0.14-11-pve. Same NIC/trigger, this time
      caught by the soft lockup watchdog with a full stack trace:
    
        watchdog: BUG: soft lockup - CPU#12 stuck for 354s! [napi/eth%d-0:329]
        CPU: 12 UID: 0 PID: 329 Comm: napi/eth%d-0 Tainted: P O L 7.0.14-11-pve #1 PREEMPT(lazy)
        RIP: 0010:atl1c_clean_tx+0x142/0x2d0 [atl1c]
        Call Trace:
         <TASK>
         __napi_poll+0x32/0x1e0
         napi_threaded_poll_loop+0x286/0x2e0
         napi_threaded_poll+0xfd/0x140
         kthread+0xf7/0x130
         ret_from_fork+0x2da/0x3a0
         ret_from_fork_asm+0x1a/0x30
         </TASK>
    
    Fixes: 43250ddd75a35d ("atl1c: Atheros L1C Gigabit Ethernet driver")
    Cc: [email protected]
    Signed-off-by: Gajdos Tamás <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: atl1e: fix soft lockup on out-of-range hw_next_to_clean read [+ + +]
Author: Gajdos Tamás <[email protected]>
Date:   Mon Sep 21 11:13:33 2026 +0200

    net: atl1e: fix soft lockup on out-of-range hw_next_to_clean read
    
    commit 374bf9e4b90f979e052332c4faca2d745c491a12 upstream.
    
    Same issue as atl1c (see the first commit in this series, "net:
    atl1c: fix soft lockup on out-of-range tpd_cons read"): the hardware
    can report an out-of-range hw_next_to_clean (seen as 0xffff) while
    the PCIe link/MAC is resetting. An out-of-range value can never be
    reached and the loop below would spin forever. Treat it as "nothing
    new to clean" instead.
    
    Fixes: a6a5325239c202 ("atl1e: Atheros L1E Gigabit Ethernet driver")
    Cc: [email protected]
    Signed-off-by: Gajdos Tamás <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: bcmgenet: do not skip WoL power up on GENET V1 [+ + +]
Author: Florian Fainelli <[email protected]>
Date:   Mon Sep 21 15:00:19 2026 -0700

    net: bcmgenet: do not skip WoL power up on GENET V1
    
    [ Upstream commit cbbc1aee7776c7fa1d89e6cb963a23e58c495dca ]
    
    bcmgenet_power_up() had an early check for bcmgenet_has_ext(priv) before
    dispatching by power mode. GENET V1 does not have the EXT block (unlike
    GENET V2+), which causes bcmgenet_power_up() to immediately return 0.
    
    As a consequence, when waking up from GENET_POWER_WOL_MAGIC on GENET V1,
    bcmgenet_wol_power_up_cfg() is never invoked to disable the WoL clock,
    clear wake event masks, and restore normal PHY and MAC operations.
    
    Move the bcmgenet_has_ext() checks to the GENET_POWER_PASSIVE and
    GENET_POWER_CABLE_SENSE cases where the EXT registers are actually
    accessed, allowing GENET_POWER_WOL_MAGIC cleanup to execute on all
    hardware versions.
    
    Fixes: c3ae64ae0c08 ("net: bcmgenet: handle GENET_POWER_WOL_MAGIC")
    Reviewed-by: Nicolai Buchwitz <[email protected]>
    Signed-off-by: Florian Fainelli <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: bcmgenet: fix 64-bit RTNL stats reading in ethtool on 32-bit systems [+ + +]
Author: Florian Fainelli <[email protected]>
Date:   Mon Sep 21 15:00:17 2026 -0700

    net: bcmgenet: fix 64-bit RTNL stats reading in ethtool on 32-bit systems
    
    [ Upstream commit 0e2bec77ea62895416600c90588f593516572bca ]
    
    When bcmgenet was converted to 64-bit statistics, STAT_RTNL members were
    switched to point into struct rtnl_link_stats64, whose fields are 64-bit
    (__u64) regardless of architecture.
    
    However, bcmgenet_get_ethtool_stats() retained a legacy check:
      if (sizeof(unsigned long) != sizeof(u32) &&
          s->stat_sizeof == sizeof(unsigned long))
    
    On 32-bit systems, sizeof(unsigned long) == sizeof(u32), causing this
    condition to evaluate to false. As a result, 64-bit RTNL stats fields were
    read via *(u32 *)p. On 32-bit Big-Endian systems (such as MIPS BE), this
    reads the high 32 bits and returns 0 until the counter exceeds 4GB; on
    32-bit Little-Endian systems (such as 32-bit ARM), the value is truncated
    to 32 bits.
    
    Fix this by checking if s->stat_sizeof == sizeof(u64) so 64-bit fields are
    always read as 64-bit values.
    
    Fixes: 59aa6e3072aa ("net: bcmgenet: switch to use 64bit statistics")
    Reviewed-by: Nicolai Buchwitz <[email protected]>
    Signed-off-by: Florian Fainelli <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: bcmgenet: initialize u64 stats seq counter for all queues [+ + +]
Author: Florian Fainelli <[email protected]>
Date:   Mon Sep 21 15:00:18 2026 -0700

    net: bcmgenet: initialize u64 stats seq counter for all queues
    
    [ Upstream commit 3aeaa609fda19c09d5298c9fedaaa3b6229601b5 ]
    
    bcmgenet_gstrings_stats statically defines ethtool statistics for queues
    0 through GENET_MAX_MQ_CNT (4). However, bcmgenet_probe() only initialized
    the u64_stats_sync seq counter up to priv->hw_params->rx_queues and
    priv->hw_params->tx_queues.
    
    Since priv->hw_params->rx_queues is 0 across all hardware versions (and
    priv->hw_params->tx_queues is 0 on GENET V1), rings 1..4 have uninitialized
    u64_stats_sync structures. When ethtool -S is run on 32-bit kernels,
    bcmgenet_get_ethtool_stats() reads stats from rx_rings[1..4], causing
    lockdep warnings due to the uninitialized sequence counters.
    
    Initialize the sequence counters for all GENET_MAX_MQ_CNT + 1 queues.
    
    Fixes: ffc2c8c4a714 ("net: bcmgenet: Initialize u64 stats seq counter")
    Reviewed-by: Nicolai Buchwitz <[email protected]>
    Signed-off-by: Florian Fainelli <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: bcmgenet: mask DMA_TIMEOUT_MASK when reading DMA_RING0_TIMEOUT [+ + +]
Author: Florian Fainelli <[email protected]>
Date:   Mon Sep 21 15:00:21 2026 -0700

    net: bcmgenet: mask DMA_TIMEOUT_MASK when reading DMA_RING0_TIMEOUT
    
    [ Upstream commit d64e277b955be4506931802837499b62c8f3968a ]
    
    bcmgenet_get_coalesce() reads DMA_RING0_TIMEOUT to calculate
    rx_coalesce_usecs without masking out bits outside DMA_TIMEOUT_MASK
    (16 bits). If upper bits are non-zero or contain status/flags, the
    computed value of rx_coalesce_usecs returned to userspace via ethtool
    becomes corrupted.
    
    Mask the register read with DMA_TIMEOUT_MASK before computing the
    timeout in microseconds.
    
    Fixes: 4a29645bfe6c ("net: bcmgenet: Implement RX coalescing control knobs")
    Reviewed-by: Nicolai Buchwitz <[email protected]>
    Signed-off-by: Florian Fainelli <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: bcmgenet: stop Tx NAPI before disabling the queues [+ + +]
Author: Nicolai Buchwitz <[email protected]>
Date:   Tue Sep 22 15:06:38 2026 +0200

    net: bcmgenet: stop Tx NAPI before disabling the queues
    
    [ Upstream commit 7e87508b5c4d81210d0a736ed01962e52f5c4c56 ]
    
    bcmgenet_netif_stop() and the Wake-on-LAN branch of bcmgenet_suspend()
    both disable the Tx queues first and stop Tx NAPI several steps later. A
    completion in flight calls netif_tx_wake_queue() in between, and nothing
    stops the queue again, so a transmit can reach the rings after they have
    been freed.
    
    Close is safe because dev_deactivate_many() stops the qdisc first.
    bcmgenet_suspend() does not, so stop Tx NAPI before the queues on both
    paths.
    
    KASAN on a Raspberry Pi CM4, driven from an MTU change because suspend
    freezes user space before the callback runs:
    
      BUG: KASAN: use-after-free in bcmgenet_xmit+0x17f8/0x2258
      Write of size 8 at addr ffffff8055844a68 by task ksoftirqd/0/14
       bcmgenet_xmit+0x17f8/0x2258
       dev_hard_start_xmit+0x13c/0x588
       sch_direct_xmit+0x108/0x340
       __dev_queue_xmit+0x1190/0x3848
    
    Fixes: 254f3239dd07 ("net: bcmgenet: revise suspend/resume")
    Signed-off-by: Nicolai Buchwitz <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: bcmgenet: validate Ethernet address in bcmgenet_set_mac_addr [+ + +]
Author: Florian Fainelli <[email protected]>
Date:   Mon Sep 21 15:00:20 2026 -0700

    net: bcmgenet: validate Ethernet address in bcmgenet_set_mac_addr
    
    [ Upstream commit 273941c85fc2632cd3e56ddff737b9245de7697d ]
    
    bcmgenet_set_mac_addr() did not check whether the provided MAC address is a
    valid Ethernet address before applying it. Userspace could configure an
    invalid address (such as all zeroes or a multicast address) while the
    interface is down.
    
    Add a call to is_valid_ether_addr() and return -EADDRNOTAVAIL if the MAC
    address is not valid.
    
    Fixes: 1c1008c793fa ("net: bcmgenet: add main driver file")
    Reviewed-by: Nicolai Buchwitz <[email protected]>
    Signed-off-by: Florian Fainelli <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: bridge: mdb: restart port group walk after deletion [+ + +]
Author: Fourie Zhang <[email protected]>
Date:   Sun Sep 20 19:08:43 2026 +0800

    net: bridge: mdb: restart port group walk after deletion
    
    commit ab1404ac81154a89fb61ac50ae9a04cd8d4834dc upstream.
    
    br_mdb_flush_pgs() keeps a pointer-to-pointer cursor while walking
    mp->ports. br_multicast_del_pg() can re-enter the same MDB entry through
    br_multicast_sg_del_exclude_ports() and unlink other port groups. If the
    cursor points into one of those groups, the next iteration dereferences a
    stale cursor and can leave mp->ports pointing at freed memory.
    
    A following RTM_GETMDB exposes the dangling pointer:
    
      BUG: KASAN: slab-use-after-free in br_mdb_dump
      Read of size 8
        br_mdb_dump
        rtnl_mdb_dump
        rtnl_dumpit
        netlink_dump
    
    Reset the cursor to mp->ports after every deletion. The deletion removes at
    least the selected group, so the restarted walk always makes progress.
    
    Fixes: a6acb535afb2 ("bridge: mdb: Add MDB bulk deletion support")
    Cc: [email protected]
    Signed-off-by: Fourie Zhang <[email protected]>
    Acked-by: Nikolay Aleksandrov <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: don't require the hwtstamp NDOs when a PHY provides timestamping [+ + +]
Author: Nicolai Buchwitz <[email protected]>
Date:   Fri Sep 18 11:55:40 2026 +0200

    net: don't require the hwtstamp NDOs when a PHY provides timestamping
    
    [ Upstream commit 31995571219c8ac30913d9c0dccad033fbb0b3da ]
    
    Removing the legacy ioctl fallback made both hwtstamp NDOs mandatory. A
    device that only timestamps in its PHY implements neither, so
    SIOCSHWTSTAMP fails with EOPNOTSUPP before anything looks at the PHY and
    PTP stops working there.
    
    The check only ever picked the legacy path. That path is gone, so drop it
    and test where the NDOs are actually called.
    
    SIOCGHWTSTAMP is new here, not restored. The old path went through
    phy_mii_ioctl(), which only handled SIOCSHWTSTAMP.
    
    Such a device now returns -ENODEV while absent instead of -EOPNOTSUPP,
    like the ones that do implement the NDOs.
    
    Fixes: 5062245a5a7f ("net: remove legacy way to get/set HW timestamp config")
    Signed-off-by: Nicolai Buchwitz <[email protected]>
    Reviewed-by: Kory Maincent <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: dsa: mt7530: fix NULL dereference on unbind of MT7531 and MT7621 [+ + +]
Author: Aleksei Sviridkin <[email protected]>
Date:   Fri Sep 18 04:50:19 2026 +0300

    net: dsa: mt7530: fix NULL dereference on unbind of MT7531 and MT7621
    
    [ Upstream commit c2cdef41e0b4d8ed23a5b41e6ad4e64594e055e4 ]
    
    The core and io supplies are only requested for ID_MT7530: both the
    devm_regulator_get() in probe and the regulator_enable() in
    mt7530_setup() are guarded by the switch id, but mt7530_remove()
    disables them unconditionally. On an MT7621 or an MT7531 both pointers
    are still NULL from devm_kzalloc(), so rmmod or a sysfs unbind calls
    regulator_disable() on NULL.
    
    Fixes: ddda1ac116c8 ("net: dsa: mt7530: support the 7530 switch on the Mediatek MT7621 SoC")
    Signed-off-by: Aleksei Sviridkin <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: dsa: mt7530: leave the MDIO IRQ mappings to regmap-irq [+ + +]
Author: Aleksei Sviridkin <[email protected]>
Date:   Fri Sep 18 04:50:20 2026 +0300

    net: dsa: mt7530: leave the MDIO IRQ mappings to regmap-irq
    
    [ Upstream commit 0d80ba0a204c6a16bd7778b50de578dff107c0fe ]
    
    mt7530_remove_common() disposes the per-PHY interrupt mappings from
    .remove, but the regmap-irq chip that owns the domain is devm-registered,
    so its parent interrupt is only freed once .remove has returned. The
    switch's own regmap-irq thread can therefore still dispatch on a mapping
    that is already gone: irq_find_mapping() returns 0, irq_to_desc() returns
    NULL and handle_nested_irq() locks desc->lock without checking it. The
    attached PHYs have not given those interrupts back yet either, which the
    kernel warns about a moment before the fault.
    
    regmap_del_irq_chip() disposes the same mappings itself, after freeing the
    parent interrupt and before removing the domain, so there is nothing left
    for the driver to do here. Until it runs the descriptors stay alive, and a
    late dispatch on one of them is harmless: dsa_unregister_switch() has freed
    the PHY handlers by then, so handle_nested_irq() finds no action and
    returns.
    
    Fixes: 254f6b272e3b ("dsa: mt7530: Utilize REGMAP_IRQ for interrupt handling")
    Signed-off-by: Aleksei Sviridkin <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: dsa: mv88e6xxx: 88E6191X and 88E6193X have no PTP [+ + +]
Author: Nicolo Giuliani <[email protected]>
Date:   Mon Sep 21 05:56:29 2026 +0200

    net: dsa: mv88e6xxx: 88E6191X and 88E6193X have no PTP
    
    [ Upstream commit 6b491af01aa5c0633a580e3b11bc7277adc903b2 ]
    
    The 88E6191X and 88E6193X are 6393 family devices that share
    mv88e6393x_ops with the 88E6393X and are marked as ptp_support. Marvell's
    UMSD driver describes both as parts without AVB (88E6193X: "BGA package -
    No AVB, No Routing, No Cut-through"), and the register access confirms it
    on an 88E6193X: the whole indirect AVB register space behind Global 2
    registers 0x16 and 0x17 reads zero, for every port, block and address,
    with the 6390 and with the 6352 command encoding. Writes to the TAI
    registers, including the clock period register and the TAI global
    configuration register, read back as zero.
    
    Since commit 7e3c18097a70 ("net: dsa: mv88e6xxx: read cycle counter
    period from hardware") the PTP setup reads the TAI clock period, so the
    switch fails to probe:
    
    mv88e6xxx ...: unexpected cycle counter period of 0 ps
    
    Add mv88e6191x_ops, a copy of mv88e6393x_ops without avb_ops and ptp_ops,
    use it for the 88E6191X and the 88E6193X and stop setting ptp_support for
    them. The 88E6393X is unchanged.
    
    Tested on an 88E6193X (Sophos XGS 107w): the switch probes and the ports
    work. I do not have an 88E6191X, it is changed because UMSD describes it
    the same way.
    
    Fixes: de776d0d316f ("net: dsa: mv88e6xxx: add support for mv88e6393x family")
    Suggested-by: Andrew Lunn <[email protected]>
    Signed-off-by: Nicolo Giuliani <[email protected]>
    Reviewed-by: Andrew Lunn <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: emac: move setting of netops to fix crash [+ + +]
Author: Christian Lamparter <[email protected]>
Date:   Mon Sep 21 18:28:17 2026 +0200

    net: emac: move setting of netops to fix crash
    
    [ Upstream commit 7c9f391ec89cb621d7af375ace2eb9a6248e5b9d ]
    
    fixes the following crash on driver initialization:
    
    |BUG: Kernel NULL pointer dereference on read at 0x00000158
    |Faulting instruction address: 0xc0566b40
    |Oops: Kernel access of bad area, sig: 11 [#1]
    |BE PAGE_SIZE=4K  PowerPC 44x Platform
    |Modules linked in:
    |CPU: 0 UID: 0 PID: 1 Comm: swapper/0 Tainted: GW 7.3.0-rc3+ #1
    |Tainted: [W]=WARN
    |Hardware name: MyBook Live APM821XX 0x12c41c83 PowerPC 44x Platform
    |NIP:  c0566b40 LR: c05648f8 CTR: c04c1e9c
    |REGS: c1053a20 TRAP: 0300   Tainted: GW  (7.3.0-rc3+)
    |MSR:  0002b000 <CE,EE,FP,ME>  CR: 24008808  XER: 00000000
    |DEAR: 00000158 ESR: 00000000
    |GPR00: c05648f8 c1053b10 c1063600 c1030000 c5ab3000 00000000 [...]
    |GPR08: 00000002 00000000 00000000 c1053b40 84002808 00000000 [...]
    |GPR16: cfffd210 00000002 c0beafcc cfffc960 00000000 c1030644 [...]
    |GPR24: c0beafbc c1037000 00000000 0000000a 00000000 c1030000 [...]
    |NIP [c0566b40] phy_link_topo_add_phy+0x2c/0x1d0
    |LR [c05648f8] phy_attach_direct+0x1a4/0x368
    |Call Trace:
    |[c1053b10] [c0811e04] klist_put+0x54/0xb4 (unreliable)
    |[c1053b40] [c05648f8] phy_attach_direct+0x1a4/0x368
    |[c1053b70] [c0564ae8] phy_connect_direct+0x2c/0x60
    |[c1053b90] [c056c7a8] of_phy_connect+0x50/0x74
    |[c1053bc0] [c0572a88] emac_probe+0xd50/0x119c
    |[c1053c90] [c04cb770] platform_probe+0x74/0xa4
    |[c1053cb0] [c04c8f08] really_probe+0x120/0x2b0
    |[c1053cd0] [c04c9254] __driver_probe_device+0x1bc/0x1fc
    |[c1053d00] [c04c9334] driver_probe_device+0x38/0xa8
    |[c1053d30] [c04c9560] __driver_attach+0xf4/0x10c
    |[c1053d50] [c04c6af8] bus_for_each_dev+0x68/0xd0
    |[c1053d90] [c04c7c34] bus_add_driver+0xcc/0x1ec
    |[c1053dc0] [c04c9f6c] driver_register+0xcc/0x110
    |[c1053de0] [c0aa3f40] emac_init+0x1c4/0x200
    
    This bug showed up starting with v7.3-rc1. At the NIP in
    phy_link_topo_add_phy() is a netdev_need_ops_lock() check.
    This was added by the following
    commit ded86da4bbb7 ("net: ethtool: relax ethnl_req_get_phydev() locking assertion")
    
    The bug shows up because at the time of_phy_connect() was called, the
    netdev_ops were *not yet* determined. My fix is to move the code that
    sets netdev_ops+commac.ops+ethtool_ops further up as emac_init_config()
    derives that by looking at the device-tree and sets the required
    dev->phy_mode accordingly.
    
    During review, the Sashiko bot's AI stated that the commac assignment
    became a dead store. Great catch! To keep the original behavior as-is,
    one mentioned option "set dev->commac.ops = &emac_commac_ops only in
    the non-gige case?" sounded like a great plan. So the dev->commac.ops
    assignment for the non-gige-case moves into the else block.
    
    This patch was tested on a WD MyBook Live (RGMII). the device now works again.
    
    Fixes: ded86da4bbb7 ("net: ethtool: relax ethnl_req_get_phydev() locking assertion")
    Signed-off-by: Christian Lamparter <[email protected]>
    Reviewed-by: Nicolai Buchwitz <[email protected]>
    Reviewed-by: Maxime Chevallier <[email protected]>
    Link: https://patch.msgid.link/49cd7343bc0e507c022071f3e2b5662b053dca73.1790007431.git.chunkeey@gmail.com
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: ena: fix MMIO read buffer leak on probe failure [+ + +]
Author: Guangshuo Li <[email protected]>
Date:   Mon Sep 21 23:42:02 2026 +0800

    net: ena: fix MMIO read buffer leak on probe failure
    
    commit 9476b4468862927297c94c440863cd8ed1e7cc83 upstream.
    
    ena_device_init() initializes the MMIO read mechanism with
    ena_com_mmio_reg_read_request_init(), which allocates a coherent DMA
    buffer for MMIO read responses.
    
    The normal removal path releases this buffer through
    ena_com_mmio_reg_read_request_destroy(). However, if ena_probe() fails
    after ena_device_init() succeeds, the error path destroys the admin
    resources and eventually frees ena_dev without destroying the MMIO read
    request, leaving the coherent DMA buffer allocated.
    
    Call ena_com_mmio_reg_read_request_destroy() in the probe error path
    before releasing the remaining device resources.
    
    This issue was found by manual code inspection.
    
    Fixes: 1738cd3ed342 ("net: ena: Add a driver for Amazon Elastic Network Adapters (ENA)")
    Cc: [email protected]
    Signed-off-by: Guangshuo Li <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: ena: fix PHC cleanup on probe failure [+ + +]
Author: Guangshuo Li <[email protected]>
Date:   Mon Sep 21 23:42:01 2026 +0800

    net: ena: fix PHC cleanup on probe failure
    
    commit 0958ea4355e2e9220ad4e13da3b7d94f365ed34f upstream.
    
    ena_probe() initializes the PHC as part of ena_device_init(), but the
    probe failure path does not destroy it before freeing the PHC private
    data.
    
    The normal removal path calls ena_phc_destroy() through
    ena_destroy_device() before ena_phc_free(). However, if probe fails
    after ena_device_init() succeeds, the error path reaches ena_phc_free()
    without unregistering the PTP clock or destroying the device PHC
    resources.
    
    Call ena_phc_destroy() in the probe error path before freeing the PHC
    private data.
    
    This issue was found by manual code inspection.
    
    Cc: [email protected] tags and describe this as a consistency cleanup
    Fixes: e0ea34158ee8 ("net: ena: Add PHC support in the ENA driver")
    Cc: [email protected]
    Signed-off-by: Guangshuo Li <[email protected]>
    Cc: stable
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: ethernet: mtk_eth_soc: unregister net_devices in case of probe failure [+ + +]
Author: Lorenzo Bianconi <[email protected]>
Date:   Wed Sep 16 15:30:13 2026 +0200

    net: ethernet: mtk_eth_soc: unregister net_devices in case of probe failure
    
    [ Upstream commit 310d1ac61a4d5a2ca8356a3a48d263acf54503ce ]
    
    If register_netdev() fails for one of the MTK_MAX_DEVS devices in
    mtk_probe(), the error path jumps to err_deinit_ppe, skipping
    mtk_unreg_dev(). The previously registered net_devices are then freed by
    mtk_free_dev() while still in NETREG_REGISTERED state, hitting the
    BUG_ON(dev->reg_state != NETREG_UNREGISTERED).
    
    Route the register_netdev() failure to err_unreg_netdev so the net_devices
    registered so far are properly unregistered before being freed.
    
    Fixes: 8a8a9e89f801 ("net: ethernet: mediatek: cleanup error path inside mtk_hw_init")
    Signed-off-by: Lorenzo Bianconi <[email protected]>
    Link: https://patch.msgid.link/20260916-mtk_eth_soc-netdev-fix-v1-1-5dac50eb65b1@oss.qualcomm.com
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails [+ + +]
Author: Coia Prant <[email protected]>
Date:   Wed Sep 23 20:37:13 2026 +0800

    net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails
    
    [ Upstream commit 8db67bb6a1fffa4df68fbbc22e39943aeeff9178 ]
    
    gmac_clk_enable() enables the bulk clocks first and then the optional
    PHY clock. If clk_prepare_enable() on the PHY clock fails, the function
    returns without rolling back the bulk clocks, and bsp_priv->clk_enabled
    stays false, so the later gmac_clk_enable(bsp_priv, false) becomes a
    no-op and the bulk clock references are leaked.
    
    Add the missing clk_bulk_disable_unprepare() on that failure path.
    
    Fixes: ea449f7fa0bf ("net: ethernet: stmmac: dwmac-rk: rework optional clock handling")
    Reviewed-by: Maxime Chevallier <[email protected]>
    Reviewed-by: Heiko Stuebner <[email protected]>
    Acked-by: Lorenzo Bianconi <[email protected]>
    Signed-off-by: Coia Prant <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: ethernet: ti: netcp: fix pm_runtime usage counter leak on error [+ + +]
Author: bui duc phuc <[email protected]>
Date:   Fri Sep 18 11:28:04 2026 +0700

    net: ethernet: ti: netcp: fix pm_runtime usage counter leak on error
    
    [ Upstream commit ac4334522e4ba4a3b6710dd5d4cc98092824b8ca ]
    
    pm_runtime_get_sync() leaves the runtime PM usage counter incremented even
    when it fails, but the error path in netcp_probe() does not call
    pm_runtime_put_noidle() to balance it, leaking a reference each time
    resume fails.
    
    Use pm_runtime_resume_and_get() instead, which automatically drops the
    usage counter on failure, fixing the leak.
    
    Fixes: 84640e27f230 ("net: netcp: Add Keystone NetCP core ethernet driver")
    Signed-off-by: bui duc phuc <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: ethtool: keep rtnl_lock for the ioctl self test [+ + +]
Author: Alexander Duyck <[email protected]>
Date:   Mon Sep 14 14:09:57 2026 -0700

    net: ethtool: keep rtnl_lock for the ioctl self test
    
    [ Upstream commit 1b82958f3f035df5ccaab5430a2302f08a5d5351 ]
    
    An offline self test that brings the interface down and back up with
    netif_close() / netif_open() requires rtnl_lock for both. Since the
    ethtool IOCTL path became rtnl-optional for ops-locked drivers, the
    ETHTOOL_TEST ioctl runs holding only the netdev instance lock, so on an
    ops-locked driver the self test now tears the device down without
    rtnl_lock.
    
    With lockdep this reproduces deterministically on every offline self
    test on such a driver; note the sole lock held is the instance lock, not
    rtnl:
    
      WARNING: suspicious RCU usage
      net/core/netpoll.c:207 suspicious rcu_dereference_protected() usage!
      1 lock held by ethtool/107:
       #0: (&dev->lock){+.+.}, at: dev_ethtool
      Call Trace:
       netpoll_poll_disable
       __dev_close_many
       netif_close_many
       netif_close
       fbnic_self_test
       dev_ethtool_locked
       dev_ethtool
       dev_ioctl
       sock_ioctl
       __x64_sys_ioctl
    
    Without lockdep the same condition trips ASSERT_RTNL() in
    __dev_close_many() / __dev_open(); that check only samples the global
    rtnl state, so it can be masked by a concurrent rtnl holder, but the
    device is still being reconfigured without the lock it requires.
    
    The ethtool self_test is a legacy ioctl-only command, so an ETHTOOL_TEST
    case is only needed on the ioctl path. Add an opt-in bit for drivers whose
    self test needs rtnl_lock and set it on the ops-locked drivers whose
    offline self test tears the interface down and up:
    
      - fbnic (ops-locked via queue_mgmt_ops): fbnic_self_test() offline path
        uses netif_close() / netif_open().
      - bnxt (ops-locked via queue_mgmt_ops): bnxt_self_test() offline path
        goes through bnxt_close_nic() / bnxt_half_open_nic() /
        bnxt_half_close_nic() / bnxt_open_nic(), which close and reopen the
        device.
    
    Fixes: f994752b1127 ("net: ethtool: optionally skip rtnl_lock on IOCTL path")
    Signed-off-by: Alexander Duyck <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/178942019771.7700.338431553546884773.stgit@ahduyck-xeon-server.home.arpa
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: flush skb_defer_nodes in dev_cpu_dead() [+ + +]
Author: Eric Dumazet <[email protected]>
Date:   Wed Sep 23 13:03:18 2026 +0000

    net: flush skb_defer_nodes in dev_cpu_dead()
    
    [ Upstream commit 06e3f54e8b22040ada01a28343badf1990b4dd7d ]
    
    When a CPU goes offline, dev_cpu_dead() drains its softnet queues
    (completion_queue, output_queue, poll_list, process_queue, and
    input_pkt_queue), but leaves net_hotdata.skb_defer_nodes untouched.
    
    If oldcpu goes offline while holding pending skbs in its
    skb_defer_nodes lists (e.g. below the sysctl_skb_defer_max >> 1 IPI
    threshold, or if the IPI races with CPU teardown), those skbs remain
    stranded until oldcpu is brought back online. If any of these skbs
    hold page_pool fragments, page_pool_destroy() will stall indefinitely
    waiting for inflight pages to be returned when a netdev or driver is
    torn down while oldcpu is offline.
    
    Additionally, if smp_call_function_single_async() fails in
    kick_defer_list_purge() because the target CPU went offline, reset
    defer_ipi_scheduled to 0 so future IPI kicks are not blocked when the
    CPU comes back online.
    
    Also, if oldcpu was the last online CPU on its NUMA node, drain that
    node's slot across all CPUs so no skbs deferred from that node remain
    stranded on idle remote CPUs (or if the node itself is subsequently
    offlined).
    
    Finally, in skb_attempt_defer_free(), re-check cpu_online(cpu) and
    whether the caller migrated CPUs after llist_add(), flushing the node
    list if so, to close the preemption TOCTOU race against CPU/node
    teardown.
    
    Fixes: 68822bdf76f1 ("net: generalize skb freeing deferral to per-cpu lists")
    Fixes: 5628f3fe3b16 ("net: add NUMA awareness to skb_attempt_defer_free()")
    Closes: https://lore.kernel.org/netdev/[email protected]/
    Signed-off-by: Eric Dumazet <[email protected]>
    Cc: Kris Pan <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: gue: reject invalid REMCSUM offsets [+ + +]
Author: Jérémy Jean <[email protected]>
Date:   Tue Sep 15 12:48:07 2026 +0000

    net: gue: reject invalid REMCSUM offsets
    
    [ Upstream commit 2566866fc30965d915d0b52b5c3323b362619f0e ]
    
    The REMCSUM option carries an absolute checksum start and checksum field
    offset. gue_remcsum() passes them to skb_remcsum_process(), whose
    partial path stores offset - start in the u16 skb->csum_offset variable.
    If offset is less than start, this underflows.
    
    A forwarded packet can retain CHECKSUM_PARTIAL and reach a NETIF_F_HW_CSUM
    driver which trusts the metadata, leading skb_copy_and_csum_dev() to write
    two bytes about 64 KiB beyond the destination buffer.
    
    Reject reversed tuples in validate_gue_flags(), after the existing length
    validation, so all GUE parsers enforce the ordering in one place.
    
    Fixes: fe881ef11cf0 ("gue: Use checksum partial with remote checksum offload")
    Signed-off-by: Jérémy Jean <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: hisilicon: hns_dsaf_mac: fix mdio device leak in hns_mac_register_phy() [+ + +]
Author: Wentao Liang <[email protected]>
Date:   Thu Sep 17 11:08:28 2026 +0000

    net: hisilicon: hns_dsaf_mac: fix mdio device leak in hns_mac_register_phy()
    
    commit 999e8295bc41d6ce45b8e54f88150efa96f3f01e upstream.
    
    hns_dsaf_find_platform_device() returns the mdio platform device with its
    reference count incremented. hns_mac_register_phy() never drops that
    reference, so the mdio device can not be released.
    
    Release the reference on both the deferred probe and the normal path.
    
    Fixes: 1d1afa2ebf82 ("net: hns: register phy device in each mac initial sequence")
    Cc: [email protected]
    Signed-off-by: Wentao Liang <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: ipconfig: bound DHCP option construction [+ + +]
Author: Yuqi Xu <[email protected]>
Date:   Sat Sep 19 16:45:27 2026 +0800

    net: ipconfig: bound DHCP option construction
    
    commit e47a1958e12abc3a17b5231a4f21c8f1bf662e08 upstream.
    
    ic_dhcp_init_options() appends the hostname (option 12), vendor-class
    (option 60) and client-ID (option 61) options into the fixed 312-byte
    bootp_pkt.exten[] buffer.  Only the client-ID branch checked the
    remaining space; the hostname and vendor-class writes were unbounded.
    
    A 64-byte hostname together with the maximum 252-byte dhcpclass=
    identifier needs 18 + (2 + 64) + (2 + 252) = 338 of the 312 available
    bytes even before the terminating END marker, so the vendor-class memcpy
    runs past the end of exten[].  With CONFIG_FORTIFY_SOURCE this is
    reported as a field-spanning write and, when the kernel is booted with
    panic_on_warn=1, aborts boot with a panic.
    
    Route the optional options through a common helper that makes sure the
    option, its 2-byte header and the END marker all fit and drops an option
    that would not.  Configurations with short options keep sending exactly
    the same bytes as before.
    
    Fixes: 130c0f47fdf9 ("ipconfig: send host-name in DHCP requests")
    Cc: [email protected]
    Reported-by: Vega <[email protected]>
    Assisted-by: LLM
    Signed-off-by: Yuqi Xu <[email protected]>
    Reviewed-by: Ren Wei <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/7808dfbfa2162dfd0b19f59aff5742d6e0db2abb.1789798023.git.xuyuqiabc@gmail.com
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: ipv6: keep room for the mac header in dst_dev_overhead() [+ + +]
Author: Yuya Kusakabe <[email protected]>
Date:   Tue Sep 22 05:49:56 2026 +0900

    net: ipv6: keep room for the mac header in dst_dev_overhead()
    
    [ Upstream commit 87cd6b717e4069dca34e0866e57eaeb44d3b173e ]
    
    The seg6, ioam6 and rpl lwtunnels size their skb_cow_head() request as
    the length they are about to push plus dst_dev_overhead(), then push the
    new headers and rebuild the mac header below them with
    skb_mac_header_rebuild().  That rebuild needs skb->mac_len of headroom,
    but dst_dev_overhead() leaves LL_RESERVED_SPACE() of the egress device,
    16 bytes for plain Ethernet.
    
    Where the mac header is longer than that, as it is on ingress through a
    VLAN device with reorder_hdr off, the rebuild runs out of room:
    skb_set_mac_header(skb, -skb->mac_len) computes a negative offset,
    stores it unchecked in the u16 skb->mac_header, and the memmove that
    follows writes skb->mac_len bytes about 64 KB past skb->head.
    Forwarding plain ping6 traffic through such a device reproduces it on
    all five seg6 encapsulation modes and on the rpl and ioam6 inline paths;
    skb->mac_header comes back as 65534 on a 704-byte head.
    
    Return the larger of the two.  The helper already returns skb->mac_len
    when it has no dst, so this only makes the other branch agree, and it
    covers every caller rather than each call site in turn.
    
    Fixes: 40475b63761a ("net: ipv6: seg6_iptunnel: mitigate 2-realloc issue")
    Fixes: dce525185bc9 ("net: ipv6: ioam6_iptunnel: mitigate 2-realloc issue")
    Fixes: 985ec6f5e623 ("net: ipv6: rpl_iptunnel: mitigate 2-realloc issue")
    Suggested-by: Andrea Mayer <[email protected]>
    Signed-off-by: Yuya Kusakabe <[email protected]>
    Reviewed-by: Justin Iurman <[email protected]>
    Reviewed-by: Gabriel Goller <[email protected]>
    Reviewed-by: Eric Dumazet <[email protected]>
    Reviewed-by: Andrea Mayer <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: libwx: fix races in Tx timestamp handling [+ + +]
Author: Jiawen Wu <[email protected]>
Date:   Mon Sep 21 15:15:49 2026 +0800

    net: libwx: fix races in Tx timestamp handling
    
    [ Upstream commit 3173cba1170131816972ed2b6185cb970ba8b747 ]
    
    wx->ptp_tx_skb is shared between the Tx path, the PTP auxiliary
    worker and the timestamp cleanup paths. The
    WX_STATE_PTP_TX_IN_PROGRESS bit prevents multiple Tx paths from
    submitting timestamp requests, but does not serialize the worker
    against cleanup.
    
    As a result, wx_ptp_clear_tx_timestamp() can free an skb after
    wx_ptp_tx_hwtstamp_work() has obtained its pointer. The worker may
    then pass the freed skb to skb_tstamp_tx() and release the same
    reference again.
    
    The cleanup path may also clear the in-progress bit while the worker
    is still processing the old skb. This allows the Tx path to publish a
    new skb which the worker can subsequently overwrite with NULL,
    leaking its reference.
    
    Add a dedicated spinlock to protect publication and consumption of
    the Tx timestamp skb. Detach the skb and clear the in-progress bit
    while holding the lock, then deliver the timestamp and release the skb
    after dropping it. Use the same locked cleanup in the quiesce path,
    but keep the detach there free of register accesses: quiesce runs
    during PCIe error recovery, where MMIO is not reliable, and it
    deliberately did not touch the device before. The lock is taken with
    interrupts disabled, because netpoll can call ndo_start_xmit() with
    hard interrupts already off.
    
    When handling a Tx DMA mapping failure, keep the transmit path
    reference until after comparing the skb under the lock. This prevents
    skb address reuse from making the error path mistake a newer timestamp
    request for the failed one.
    
    Fixes: 06e75161b9d4 ("net: wangxun: Add support for PTP clock")
    Reported-by: Sashiko <[email protected]>
    Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/6C7EC12D69217315%2B20260818074721.45536-1-jiawenwu%40trustnetic.com
    Signed-off-by: Jiawen Wu <[email protected]>
    Link: https://patch.msgid.link/77431AF9A0E369F3+20260921071549.1141804-1-jiawenwu@trustnetic.com
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: macb: fix dma_alloc_coherent() leak on macb_alloc() error paths [+ + +]
Author: Théo Lebrun <[email protected]>
Date:   Fri Sep 18 21:53:52 2026 +0200

    net: macb: fix dma_alloc_coherent() leak on macb_alloc() error paths
    
    commit 23d42b9a3bcd55b17d3b371544fedc708c2397e9 upstream.
    
    Fix 3 leaks in macb_alloc() error paths:
    - Tx buffer allocated but crossing a 4G boundary: Tx leaked.
    - Rx buffer allocation fails: Tx leaked.
    - Rx buffer allocated but crossing a 4G boundary: Tx & Rx leaked.
    
    This is because our error handling calls macb_free(bp) which in turn
    frees the buffers stored in bp->queues[0], but nothing has been stored
    in there. Fix by storing allocated buffers into bp->queues[0] ASAP.
    
    Fixes: 78d901897b3c ("net: macb: single dma_alloc_coherent() for DMA descriptors")
    Cc: [email protected]
    Signed-off-by: Théo Lebrun <[email protected]>
    Reviewed-by: Nicolai Buchwitz <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: mdio: realtek-rtl9300: fix RTL931x C22 extended page selection [+ + +]
Author: Jonas Jelonek <[email protected]>
Date:   Fri Sep 18 21:19:55 2026 +0000

    net: mdio: realtek-rtl9300: fix RTL931x C22 extended page selection
    
    [ Upstream commit 89a8a1eef2d441b7825a6c0116ce817235a7ffe6 ]
    
    The RTL931x indirect access engine has a separate nine-bit extended page
    field. The driver leaves it at zero, and otto_emdio_run_cmd() therefore
    programs extended page zero for every Clause 22 transaction. This
    overrides page selection made through PHY register 30, causing accesses
    to private PHY pages to hit extended page zero instead.
    
    Set the field to its 0x1ff "do not change" value for RTL931x Clause 22
    reads and writes. This preserves extended page selection made through
    PHY register 30 and restores access to its private register pages.
    
    Fixes: 5ebdcac59aff ("net: mdio: realtek-rtl9300: Add support for RTL931x")
    Signed-off-by: Jonas Jelonek <[email protected]>
    Acked-by: Markus Stockhausen <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: openvswitch: conntrack: avoid modifying shared unconfirmed ct entry [+ + +]
Author: Ilya Maximets <[email protected]>
Date:   Mon Sep 21 16:55:43 2026 +0200

    net: openvswitch: conntrack: avoid modifying shared unconfirmed ct entry
    
    commit 26b2bd70d22457556e2fa01cbf1192cb1a94d619 upstream.
    
    In a case where skb with an unconfirmed ct entry gets cloned, we may
    end up committing both but with different sets of extensions.
    
    The series of events:
    
     1. The first clone wants to commit and runs the helpers wiring up
        the extension pointer into the expectation list.
     2. Then it looses the confirmation keeping the entry unconfirmed.
     3. Second clone now wants to commit labels and adds the new extension
        for that breaking the pointer in the expectation list causing
        UAF on the destruction path later.
    
    While this is possible to trigger, there should be no practical
    network pipeline where committing both clones without modifications
    into the same zone is needed.  So, let's just reset the entry in case
    for some reason we got an skb with a shared one during commit.  This
    doesn't affect any known use cases, but avoids any potential problems
    with sharing and modification of the unconfirmed ct entry.
    
    The fixes tag points to the introduction of helpers, since that's the
    main UAF trigger for the sharing.
    
    Fixes: cae3a2627520 ("openvswitch: Allow attaching helpers to ct action")
    Cc: [email protected]
    Reported-by: Axel Mierczuk <[email protected]>
    Signed-off-by: Ilya Maximets <[email protected]>
    Reviewed-by: Aaron Conole <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: openvswitch: conntrack: fix helper UAF due to extensions realloc [+ + +]
Author: Ilya Maximets <[email protected]>
Date:   Mon Sep 21 16:55:45 2026 +0200

    net: openvswitch: conntrack: fix helper UAF due to extensions realloc
    
    commit 1a4151e6be57b098b7a5ebfbde58585e83200cdc upstream.
    
    While calling the helpers, a raw pointer to the extensions area is
    wired into expectations list:
    
      -> nf_ct_helper()
       -> helper->help()
        -> nf_ct_expect_related_report()
         -> nf_ct_expect_insert()
          -> hlist_add_head_rcu(&exp->lnode, &master_help->expectations)
    
    In case the connection is not confirmed yet, more extensions can be
    added afterwards with *_ext_add() calls reallocating the extension
    space and leaving the now invalid pointer in the expectations list
    that is later accessed while removing the expectation.
    
    Make sure that helpers are called at the end after all the other
    extensions are already added.
    
    Note that the helper rejection now leaves the mark and labels set,
    but that's not different from how the NAT was handled before or how
    the mark and the labels were handled on confirmation failure.  And
    there are no atomicity guarantees provided by the API anyway.
    
    Fixes: cae3a2627520 ("openvswitch: Allow attaching helpers to ct action")
    Cc: [email protected]
    Reported-by: Axel Mierczuk <[email protected]>
    Signed-off-by: Ilya Maximets <[email protected]>
    Reviewed-by: Aaron Conole <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: openvswitch: conntrack: remove 'add_helper' dead code [+ + +]
Author: Ilya Maximets <[email protected]>
Date:   Mon Sep 21 16:55:44 2026 +0200

    net: openvswitch: conntrack: remove 'add_helper' dead code
    
    commit 5e6c14dd42a1c1fe938e573dc6c9098145b2b0c4 upstream.
    
    This variable can only become 'true' when the connection is not
    confirmed, but it is only checked when it is confirmed.  So, it can be
    treated as being always false and just removed.
    
    Fixes: 3c1860543fcc ("openvswitch: add nf_ct_is_confirmed check before assigning the helper")
    Cc: [email protected]
    Signed-off-by: Ilya Maximets <[email protected]>
    Reviewed-by: Aaron Conole <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: pcs: rzn1-miic: Fix config array initialization [+ + +]
Author: Kyle Hendry <[email protected]>
Date:   Tue Sep 15 10:39:20 2026 -0700

    net: pcs: rzn1-miic: Fix config array initialization
    
    [ Upstream commit daf677c2c6449011ee695d55b48b5b2977a36f88 ]
    
    Fix memset parameters to initialize the entire DT value array
    
    Fixes: f39e968dc168a7bd ("net: pcs: rzn1-miic: Move configuration data to SoC-specific struct")
    Reviewed-by: Geert Uytterhoeven <[email protected]>
    Signed-off-by: Kyle Hendry <[email protected]>
    Reviewed-by: Lad Prabhakar <[email protected]>
    Link: https://patch.msgid.link/20260915-rzn1-miic-fix-array-v5-1-b7173fd5b97d@reliablecontrols.com
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: pcs: xpcs: fix clock reference leak on xpcs_init_clks failure [+ + +]
Author: Coia Prant <[email protected]>
Date:   Sun Sep 20 01:20:21 2026 +0800

    net: pcs: xpcs: fix clock reference leak on xpcs_init_clks failure
    
    [ Upstream commit 9892d71cf0ce3ff3d4fed2d9a3968fd4feb1c918 ]
    
    xpcs_init_clks() takes references with clk_bulk_get_optional() and then
    enables them with clk_bulk_prepare_enable(). If the enable step fails,
    the function returns without dropping the references.
    
    xpcs_create() handles the failure through out_free_data, which calls
    xpcs_free_data() but never xpcs_clear_clks(), so the clk references are
    leaked.
    
    Add the missing clk_bulk_put() on the enable failure path. The
    prepare/enable side is already rolled back by
    clk_bulk_prepare_enable() itself.
    
    Fixes: f6bb3e9d98c2 ("net: pcs: xpcs: Add Synopsys DW xPCS platform device driver")
    Signed-off-by: Coia Prant <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: phy: intel-xway: workaround 100BASE-TX Link-Up issue [+ + +]
Author: Alexander Sverdlin <[email protected]>
Date:   Tue Sep 22 09:52:46 2026 +0200

    net: phy: intel-xway: workaround 100BASE-TX Link-Up issue
    
    commit b94773dc4df7026a6f29b2c65e96e88d29cdb576 upstream.
    
    MaxLinear GSW12x/GSW14x Ethernet Switch Errata Sheet states:
    "An issue has been sporadically observed after device power-on on the first
    link-up attempt in 100BASE-TX mode resulting in either the link-up taking a
    long time, or failing to link-up altogether...
    
    Workaround:
    After power-on, enable Cable Diagnostic Mode for all ports and disable
    it..."
    
    Implement the proposed workaround unconditionally in the Intel XWAY driver
    (MaxLinear GSW1xx switches incorporate Intel XWAY PHYs) because the
    diagnostic bits have the same meaning even in older integral PHYs such as
    GPY111/PEF7071/PHY11G. So it's not clear how to distinguish the affected
    newer integrated PHYs, but the workaround should not hurt the older PHYs.
    
    Cc: [email protected]
    Fixes: 22335939ec90 ("net: dsa: add driver for MaxLinear GSW1xx switch family")
    Signed-off-by: Alexander Sverdlin <[email protected]>
    Reviewed-by: Andrew Lunn <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: phy: micrel: Advance register data pointer in write loop [+ + +]
Author: Abhishek Ojha <[email protected]>
Date:   Wed Sep 16 19:19:28 2026 -0400

    net: phy: micrel: Advance register data pointer in write loop
    
    commit 95c4d54ed02283e9a09e8cd7360e384daa67a741 upstream.
    
    lanphy_write_reg_data() does not advance the data pointer while iterating
    over the register table. As a result, it writes the first entry num times
    and leaves the remaining errata registers unconfigured.
    
    Single-entry tables are unaffected, but tables with multiple entries
    leave every entry after the first unapplied.
    
    Advance the data pointer after each successful write so every table entry
    is applied in order.
    
    Fixes: c8732e933925 ("net: phy: micrel: lan8842 errata")
    Cc: [email protected]
    Signed-off-by: Abhishek Ojha <[email protected]>
    Reviewed-by: Andrew Lunn <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: phylink: record the PHY only once bringup cannot fail [+ + +]
Author: Aleksei Sviridkin <[email protected]>
Date:   Mon Sep 21 01:20:44 2026 +0300

    net: phylink: record the PHY only once bringup cannot fail
    
    [ Upstream commit a940003f44e7e441c228151dd212642152700ec8 ]
    
    phylink_bringup_phy() stores the PHY in pl->phydev before its last
    fallible step: on a MAC whose phylink ops implement LPI,
    phy_eee_rx_clock_stop() can fail with a real MDIO error. The callers
    unwind with phy_detach(), which knows nothing about pl->phydev, so a
    pointer to a PHY that is no longer attached outlives the failed
    connect.
    
    What that costs depends on how the caller got here.
    phylink_connect_phy() goes through phylink_attach_phy(), which refuses
    to attach while pl->phydev is set, turning a transient MDIO error into
    a permanent -EBUSY. The SFP path is worse than that: sfp_sm_probe_phy()
    answers the failure with phy_device_remove() and phy_device_free(), and
    it assigns sfp->mod_phy only past that error return, so nothing clears
    pl->phydev and it is left pointing at a freed phy_device that
    phylink_resolve() and the ethtool helpers go on reading.
    phylink_fwnode_phy_connect() has no such check, so a later connect
    overwrites the stale pointer and hides the problem. A disconnect does
    not: phylink_disconnect_phy() hands that pointer to phy_disconnect(),
    and the second phy_detach() on the same PHY drops references the first
    one already released.
    
    Found while making a DSA port survive a PHY whose driver arrives after
    the switch probes: keeping the port across a failed connect and
    retrying is what makes this window reachable.
    
    Publish the pointer after the last call that can fail instead of
    unwinding it afterwards. Nothing between the two points reads
    pl->phydev, and the registration that follows cannot fail:
    phy_request_interrupt() falls back to polling on its own. The PHY-side
    state keeps the order it had, so no MDIO operation moves relative to
    another.
    
    Fixes: 03abf2a7c654 ("net: phylink: add EEE management")
    Signed-off-by: Aleksei Sviridkin <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: skbuff: fix pull-bound underflow in skb_checksum_setup_ipv6() [+ + +]
Author: Shihuang Liu <[email protected]>
Date:   Sat Sep 19 21:36:04 2026 +0800

    net: skbuff: fix pull-bound underflow in skb_checksum_setup_ipv6()
    
    [ Upstream commit 3b4e0b0c008a8c1b474730248cd5b873026c74bd ]
    
    skb_maybe_pull_tail() subtracts skb_headlen(skb) from the unsigned max
    argument and passes the result to __pskb_pull_tail() as a signed int.  The
    function does not ensure that max is at least skb_headlen(skb).
    
    This can happen while parsing IPv6 extension headers when an skb already
    has a linear area larger than MAX_IPV6_HDR_LEN.  Once the parser needs data
    beyond the linear area, max - skb_headlen(skb) wraps and is converted to a
    negative delta.  __pskb_pull_tail() then passes that negative length to
    skb_copy_bits(), where it can become a very large copy length.
    
    Pass the requested length itself as the pull bound at the three
    extension-header call sites, so the delta can no longer go negative.
    
    Fixes: 1431fb31ecba ("xen-netback: fix fragment detection in checksum setup")
    Suggested-by: Eric Dumazet <[email protected]>
    Signed-off-by: Shihuang Liu <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: spacemit: clear TX descriptor on fragment mapping failure [+ + +]
Author: Muhammad Bilal <[email protected]>
Date:   Sun Sep 20 00:19:37 2026 +0500

    net: spacemit: clear TX descriptor on fragment mapping failure
    
    [ Upstream commit 2d14720beb58870b15a52b236c3ab0be0e06e915 ]
    
    emac_tx_mem_map() writes TX_DESC_0_OWN into the ring descriptor for
    every slot beyond old_head as soon as that slot's memset()'d local
    copy is committed with "*tx_desc_addr = tx_desc", i.e. before the
    buffers for that slot have necessarily all been mapped successfully.
    If emac_tx_map_frag() then fails on a later fragment, the err_free_skb
    path calls emac_free_tx_buf() to unmap and drop the skb, but leaves
    the already-written descriptor memory untouched, and tx_ring->head is
    never advanced past old_head (the "tx_ring->head = head" store is
    skipped by the goto).
    
    So a slot between old_head and the rolled-back head can be left with
    TX_DESC_0_OWN set and buffer_addr_{1,2} pointing at DMA mappings that
    emac_free_tx_buf() just tore down, while software considers that slot
    free again. The next successful emac_tx_mem_map() call only rebuilds
    old_head itself; if the DMA engine auto-advances into the following
    descriptor once it finishes old_head's packet, it will fetch that
    stale, already-unmapped address.
    
    emac_tx_clean_desc() already treats emac_free_tx_buf() and clearing
    the descriptor as a pair when reclaiming completed descriptors; do
    the same in the mapping failure path.
    
    Fixes: bfec6d7f2001 ("net: spacemit: Add K1 Ethernet MAC")
    Signed-off-by: Muhammad Bilal <[email protected]>
    Reviewed-by: Vivian Wang <[email protected]>
    Reviewed-by: Troy Mitchell <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: stmmac: clear stale buf->page after recycling on skb build failure [+ + +]
Author: Lorenzo Bianconi <[email protected]>
Date:   Mon Sep 21 16:46:18 2026 +0200

    net: stmmac: clear stale buf->page after recycling on skb build failure
    
    [ Upstream commit 0a7822e34a0bfde31b194ac3da3253e5032b44cc ]
    
    In stmmac_rx(), when napi_build_skb() fails the descriptor page is
    recycled back to the page pool with page_pool_recycle_direct(), but
    buf->page is left pointing at the recycled page, unlike every other
    consumption site in the function which clears the pointer after handing
    the page away.
    
    With the stale pointer stmmac_rx_refill() skips the replacement
    allocation and programs the already-recycled page back into the RX
    descriptor.
    
    Clear buf->page on the napi_build_skb() failure path to keep the buffer
    lifecycle consistent with the other consumption sites.
    
    Fixes: df542f669307 ("net: stmmac: Switch to zero-copy in non-XDP RX path")
    Signed-off-by: Lorenzo Bianconi <[email protected]>
    Reviewed-by: Maxime Chevallier <[email protected]>
    Link: https://patch.msgid.link/20260921-stmmac-fix-napi-build-skb-error-v1-1-3d54bf6d9bb6@oss.qualcomm.com
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: stmmac: dwmac4: Use the correct bufzise when the len is exactly 8K [+ + +]
Author: Maxime Chevallier <[email protected]>
Date:   Thu Sep 17 23:53:36 2026 +0200

    net: stmmac: dwmac4: Use the correct bufzise when the len is exactly 8K
    
    [ Upstream commit b42e7012773a0e81e97a2dda6ef907f5147a6658 ]
    
    DMA bufsize selection isn't made on the MTU but the actual frame length,
    so including the L2 header. On DWMAC4, if the len is exactly BUF_SIZE_8KiB,
    the next larger size is incorrectly selected.
    
    Lets fix the comparison and while at it, rename the parameter from len
    to mtu.
    
    Fixes: c3efed5ad1b0 ("net: stmmac: Enable dwmac4 jumbo frame more than 8KiB").
    Signed-off-by: Maxime Chevallier <[email protected]>
    Reviewed-by: Nicolai Buchwitz <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: stmmac: selftests: Account for alignment shift on dwmac1000 for Jumbo test [+ + +]
Author: Maxime Chevallier <[email protected]>
Date:   Thu Sep 17 23:53:38 2026 +0200

    net: stmmac: selftests: Account for alignment shift on dwmac1000 for Jumbo test
    
    [ Upstream commit c4ac6e94eb9423126bda907a7f2933284f7450ff ]
    
    On dwmac1000, we currently only support single-descriptor frames. The
    Jumbo test started failing when NET_IP_ALIGN was added to align the IP
    header, as this tests tries to send the biggest possible frame.
    
    On dwmac1000 the DMA transfer is aligned on 4-bytes, so adding a 2-byte
    shift at the start-of-buffer address means it takes a whole extra 4-byte
    DMA burst to receive the Jumbo packet, causing it to spill over the next
    descriptor.
    
    This doesn't seem to happen on dwmac4 and xgmac that appear to correctly
    handle unaligned xfers (only tested on dwmac4)
    
    Let's account for that in the Jumbo test, reduce the size of our big
    packet by the align size.
    
    Fixes: 23680bf5f8c6 ("net: stmmac: restore NET_IP_ALIGN in the RX DMA offset")
    Reviewed-by: Nicolai Buchwitz <[email protected]>
    Signed-off-by: Maxime Chevallier <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: stmmac: selftests: Capture all packets for vlan checks [+ + +]
Author: Maxime Chevallier <[email protected]>
Date:   Thu Sep 17 23:53:35 2026 +0200

    net: stmmac: selftests: Capture all packets for vlan checks
    
    [ Upstream commit 960db6f65788c21249ea04a947d5c01e38d19294 ]
    
    While we use vlan_vid_add to trigger the tag filtering machinery
    in the driver, there's no netdev associated to the VLAN. This causes the
    skb to arrive with empty skb->vlan_tci fields, as the packet is marked
    OTHERHOST in __netif_receive_skb_core(), and we fail our validation.
    
    Let's use the proxy mechanism introduced for DSA, that registers a
    ETH_P_ALL packet handler that runs earlier, before the vlan netdev
    lookup, then filters for the correct ethertype before passing an skb
    clone to our validation function.
    
    As we may receive external frames with the right tag from the outside,
    let's move the address check in the vlan validation function earlier.
    
    Fixes: 091810dbded9 ("net: stmmac: Introduce selftests support")
    Reviewed-by: Nicolai Buchwitz <[email protected]>
    Signed-off-by: Maxime Chevallier <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: stmmac: selftests: Check the dev->features for S-TAG offload testing [+ + +]
Author: Maxime Chevallier <[email protected]>
Date:   Thu Sep 17 23:53:34 2026 +0200

    net: stmmac: selftests: Check the dev->features for S-TAG offload testing
    
    [ Upstream commit ba804b23d76d278ee475b8427fa7c5623ce5e270 ]
    
    The S-TAG offload insertion incorrectly checks the dvlan (double vlan)
    DMA cap, which is different than S-TAG support. Use
    NETIF_F_HW_VLAN_STAG_TX to check if the feature is supported instead.
    
    Note that this flag isn't set in stmmac yet, but contrary to ARP
    offload, this is a feature that has a chance to get there eventually so
    let's leave the selftest here for now. It'll report -EOPNOTSUPP in the
    meantime.
    
    Fixes: 091810dbded9 ("net: stmmac: Introduce selftests support")
    Reviewed-by: Nicolai Buchwitz <[email protected]>
    Signed-off-by: Maxime Chevallier <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: stmmac: selftests: Support running selftests on DSA conduits [+ + +]
Author: Maxime Chevallier <[email protected]>
Date:   Thu Sep 17 23:53:32 2026 +0200

    net: stmmac: selftests: Support running selftests on DSA conduits
    
    [ Upstream commit d68acbf93531abdb5b02b21994cd4c15a3c95b42 ]
    
    Most stmmac selftests rely on dev_add_pack() to add custom handlers,
    that validate the packets sent to ourselves through MAC loopback.
    
    However, when the stmmac-driven interface is a DSA CPU conduit, all
    frames that are received have ETH_P_XDSA as a protocol, even though they
    don't actually contain any tag as they come from the loopback and not
    the switch.
    
    This will prevent any incoming packet to match our packet handlers.
    
    Let's register a ETH_P_ALL packet handler when we detect that we're a
    DSA conduit, and use a proxy packet handler to filter the h_proto.
    
    As this allows external frames to be received through our .func(), the
    packet handler is added after the dev->addr field is populated in our
    selftest attributes.
    
    Note that we may still receive incoming packets from the switch, but
    these frames shouldn't interfere with the very specific frames used for
    selftests, and stmmac selftests in general aren't safe against external
    traffic interferences.
    
    This was validated on a WPQ864 devkit for IPQ8064, that has the SoC
    connected to a QCA8k switch.
    
    The ARP offload's packet handler is left alone, this feature is just not
    implemented in stmmac and due for removal.
    
    Fixes: 091810dbded9 ("net: stmmac: Introduce selftests support")
    Reviewed-by: Nicolai Buchwitz <[email protected]>
    Signed-off-by: Maxime Chevallier <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: stmmac: selftests: Validate EEE based on the actual LPI timer value [+ + +]
Author: Maxime Chevallier <[email protected]>
Date:   Thu Sep 17 23:53:33 2026 +0200

    net: stmmac: selftests: Validate EEE based on the actual LPI timer value
    
    [ Upstream commit c8c1795aa8106293020a836d01e127e98f442925 ]
    
    The EEE selftest is a 2-step test :
     - It validates that we enter in LPI mode with the
       irq_tx_path_in_lpi_mode_n counter
     - It then validates that we exit LPI when sending a frame, with the
       irq_tx_path_exit_lpi_mode_n counter.
    
    The current state of the test lacks 2 main things :
    
     - We don't know exactly when was the previous frame sent (it's from the
       previous selftest)
    
     - The timeout is hardcoded, while the LPI is entered after a
       user-configurable delay. On top of that, the timeout loop uses a
       pre-decrement iterator (--retries) that actually only iterate nine
       times, so 900ms while the default LPI value is 1 second.
    
    Let's therefore make it more deterministic :
    
     - Send a frame at the beginning of the test
     - Wait for more than the lpi timer value, we timeout after about twice
       the value,
     - Then send another frame, and verify that we do go out of LPI, also
       with a timeout.
    
    As LPI timer can get pretty high, bail out if LPI timer is over 5
    seconds.
    
    Note that the test's goal isn't to validate the LPI timer value itself,
    only that we enter/leave LPI mode.
    
    Fixes: 091810dbded9 ("net: stmmac: Introduce selftests support")
    Reviewed-by: Nicolai Buchwitz <[email protected]>
    Signed-off-by: Maxime Chevallier <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: stmmac: size the RX buffers from the frame length, not the MTU [+ + +]
Author: Maxime Chevallier <[email protected]>
Date:   Thu Sep 17 23:53:37 2026 +0200

    net: stmmac: size the RX buffers from the frame length, not the MTU
    
    [ Upstream commit b8a26d46c0a4254f8bfe143681adb9d18d020299 ]
    
    When picking the buffsize to use based on the MTU, we shouldn't check
    only the MTU value, but also :
     - ETH_HLEN for the L2 header,
     - up to 2 VLAN tags,
     - the FCS,
    
    The default bufsize is 1536 bytes, which is enough to contain all the
    above so this hasn't surfaced before, but the addition of NET_IP_ALIGN
    to the start of buffer address tripped the Jumbo selftest, leading to
    this discovery.
    
    With that, we don't need the '>=' checks on the buffer len, we can use
    more consistent comparison operators in stmmac_set_bfsize.
    
    Fixes: 286a83721720 ("stmmac: add CHAINED descriptor mode support (V4)")
    Reviewed-by: Nicolai Buchwitz <[email protected]>
    Signed-off-by: Maxime Chevallier <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: txgbe: fix FDIR filter restore for VF rules [+ + +]
Author: Zhang Yunfei <[email protected]>
Date:   Fri Sep 11 17:11:23 2026 +0800

    net: txgbe: fix FDIR filter restore for VF rules
    
    commit 651010592bdce7005c1179498327e51bfc4fe1a5 upstream.
    
    txgbe_fdir_filter_restore() reprograms every filter from
    txgbe->fdir_filter_list after a reset. It extracts the ring part of
    filter->action with ethtool_get_flow_spec_ring() and maps it onto a
    PF rx ring, silently dropping the VF part of the cookie that
    txgbe_add_ethtool_fdir_entry() stores there (input->action =
    fsp->ring_cookie).
    
    For a rule directed at a VF, restore therefore reprograms the filter
    to the PF queue with the same ring index: after any down/up or
    txgbe_reinit_locked(), traffic matching the rule is steered to the
    PF instead of the VF.
    
    Handle VF rules the same way txgbe_add_ethtool_fdir_entry() does:
    validate vf against wx->num_vfs and ring against
    wx->num_rx_queues_per_pool, and map the ring onto the absolute
    queue index ((vf - 1) * wx->num_rx_queues_per_pool) + ring.
    
    Fixes: 7a91722e0dd4 ("net: txgbe: Support the FDIR rules assigned to VFs")
    Cc: [email protected]
    Signed-off-by: Zhang Yunfei <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: usb: catc: bound the RX packet length in catc_rx_done() [+ + +]
Author: Aamir Ahmed <[email protected]>
Date:   Tue Sep 15 00:06:58 2026 +0100

    net: usb: catc: bound the RX packet length in catc_rx_done()
    
    [ Upstream commit 9d565b6b72fe3f41fd43636e143072848105189f ]
    
    catc_rx_done() walks a multi-packet URB, reading a two-byte length from
    each packet header. Its bound, pkt_len > urb->actual_length, ignores the
    header offset and compares against the whole transfer rather than the
    bytes left from pkt_start, so a crafted packet header makes
    skb_copy_to_linear_data() read past the buffer.
    
    A length below ETH_HLEN is also accepted, including zero, and
    eth_type_trans() then reads a MAC header from the uninitialised tailroom
    of a shorter skb. The is_f5u011 branch takes its length straight from
    the transfer, so a zero-length URB reaches the same path.
    
    Track the bytes remaining from the current packet, and reject a header
    that does not fit, a length past what is left, and a length below an
    Ethernet header.
    
    A transfer shorter than an Ethernet header, including a zero-length one,
    previously became a runt skb passed to netif_rx() and counted as
    received; it is now counted in rx_length_errors and ends the walk.
    
    Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
    Signed-off-by: Aamir Ahmed <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/AS8P251MB00015FD7716F38C345619B56C8BB2@AS8P251MB0001.EURP251.PROD.OUTLOOK.COM
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: usb: cdc_mbim: add MeiG Smart SRM821 to ZLP whitelist [+ + +]
Author: Ming Wang <[email protected]>
Date:   Sun Sep 20 15:44:59 2026 +0800

    net: usb: cdc_mbim: add MeiG Smart SRM821 to ZLP whitelist
    
    commit f75f21ef36285e5f56ee0c428bd2909ee81165b9 upstream.
    
    The MeiG Smart SRM821 5G module (0x2dee:0x4d53) crashes and drops off
    the USB bus when it receives a Zero Length Packet (ZLP) after sending
    or receiving an NTB of exactly 16384 bytes (tx_max).
    
    According to the MBIM specification, devices do not require a ZLP
    if the NTB size is exactly dwNtbOutMaxSize. However, the cdc_mbim
    driver defaults to sending ZLPs for devices not explicitly whitelisted
    to accommodate non-conformant hardware. This default behavior breaks
    the strictly conformant MeiG SRM821 module.
    
    Add this device to the ZLP conformance whitelist (cdc_mbim_info) so
    the driver will pad the NTB to avoid sending ZLPs, preventing the
    device firmware from crashing.
    
    Cc: [email protected]
    Signed-off-by: Ming Wang <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: usb: lan78xx: Fix URB reference leak in lan78xx_submit_deferred_urbs() [+ + +]
Author: Wentao Liang <[email protected]>
Date:   Thu Sep 17 11:58:11 2026 +0000

    net: usb: lan78xx: Fix URB reference leak in lan78xx_submit_deferred_urbs()
    
    commit 17741334d00bf5ebd37f8c1c36bc9c146a351deb upstream.
    
    usb_get_from_anchor() hands over a reference to the URB, which the caller
    must release. lan78xx_submit_deferred_urbs() never does, so every deferred
    Tx URB keeps an extra reference: the counter grows on each suspend/resume
    cycle and the URBs are never freed when the buffers are released. Drop
    the reference after submitting, and on the path that drops the packet
    instead of submitting it.
    
    Fixes: 5f4cc6e25148 ("lan78xx: Fix race conditions in suspend/resume handling")
    Cc: [email protected]
    Signed-off-by: Wentao Liang <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

net: usb: sr9700: include receive overhead in the length check [+ + +]
Author: Pengpeng Hou <[email protected]>
Date:   Sun Sep 20 11:47:45 2026 +0800

    net: usb: sr9700: include receive overhead in the length check
    
    [ Upstream commit c06bde80ae7a7b595732f7cabcb92cf08db9d56a ]
    
    The receive fixup subtracts the Ethernet CRC from the reported packet
    length, but compares that payload length against the whole remaining
    receive buffer. The following copy starts after the three-byte header,
    and the cursor advance consumes both that header and the four-byte CRC.
    
    Require the payload to fit after SR_RX_OVERHEAD before copying it or
    advancing to the next packet. The loop already ensures that the
    remaining buffer is larger than the overhead, so the subtraction is
    safe.
    
    The issue was found by our static-analysis tool.
    
    Fixes: c9b37458e956 ("USB2NET : SR9700 : One chip USB 1.1 USB2NET SR9700Device Driver Support")
    Reviewed-by: Ethan Nelson-Moore <[email protected]>
    Tested-by: Ethan Nelson-Moore <[email protected]>
    Signed-off-by: Pengpeng Hou <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

net: wangxun: implement soft quiesce for PCIe error recovery [+ + +]
Author: Jiawen Wu <[email protected]>
Date:   Mon Aug 3 14:43:33 2026 +0800

    net: wangxun: implement soft quiesce for PCIe error recovery
    
    [ Upstream commit c023e9769de94cb7b7897297e71f50b0f436c473 ]
    
    Function wx_soft_quiesce() provide a lightweight shutdown path during
    PCIe error recovery. It avoids MMIO-dependent operations in PCIe error
    status.
    
    Waiting for the service task to complete may unnecessarily delay PCIe
    error recovery, especially if the work item is already blocked by the
    hardware failure that triggered AER. So the service task is not
    explicitly cancelled in quiesce path. As a measure to block the service
    task, the checking of WX_STATE_DOWN and WX_STATE_RESETTING is added at
    the entry of relevant work item.
    
    Signed-off-by: Jiawen Wu <[email protected]>
    Reviewed-by: Aleksandr Loktionov <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Stable-dep-of: 3173cba11701 ("net: libwx: fix races in Tx timestamp handling")
    Signed-off-by: Sasha Levin <[email protected]>

net: xps: reject an out of range traffic class [+ + +]
Author: Norbert Szetei <[email protected]>
Date:   Mon Sep 21 17:03:57 2026 +0200

    net: xps: reject an out of range traffic class
    
    [ Upstream commit 4da3b7b8b50f3e2fde54a4c18a82a8e3f6223910 ]
    
    Only the entries below dev->num_tc are valid in dev->tc_to_txq[], and
    dev->prio_tc_map[] may only name classes below it. netdev_set_num_tc()
    lowers dev->num_tc without touching either array.
    
    netdev_txq_to_tc() walks all TC_MAX_QUEUE slots and
    netdev_get_prio_tc_map() returns the entry as it stands, so a leftover
    entry is handed out as a traffic class >= dev->num_tc. Taking that
    class from netdev_txq_to_tc(), __netif_set_xps_queue() rejects only a
    negative one and indexes an XPS map sized for dev->num_tc classes:
    
            tci = j * num_tc + tc;
            RCU_INIT_POINTER(new_dev_maps->attr_map[tci], map);
    
    attr_map[] holds nr_ids * num_tc entries and j runs over the ids named
    in the mask, so a class that is not below num_tc pushes tci past the end
    of the map for the last ids and the store overruns it.
    
    Any caller that lowers num_tc leaves such entries behind, and
    mqprio_destroy() tears down with netdev_set_num_tc(dev, 0) rather than
    netdev_reset_tc(). After mqprio with 8 classes then 1, tc_to_txq[1..7]
    still describe txq 1..7. The splat is from an XPS write to txq 2 on a
    veth with 8 rx queues: attr_map[] has 8 * 1 entries, tci = j + 2, and
    j == 6 stores one past the end of the 88-byte map:
    
      BUG: KASAN: slab-out-of-bounds in __netif_set_xps_queue (net/core/dev.c:2954)
      Write of size 8 at addr ffff88813016bc58 by task xps_oob/634
       __netif_set_xps_queue (net/core/dev.c:2954)
       xps_rxqs_store (net/core/net-sysfs.c:1880)
       netdev_queue_attr_store (net/core/net-sysfs.c:1390)
      Allocated by task 634:
       __kmalloc_noprof (mm/slub.c:5439)
       __netif_set_xps_queue (net/core/dev.c:2937)
      The buggy address is located 0 bytes to the right of
       allocated 88-byte region [ffff88813016bc00, ffff88813016bc58)
    
    Reject a class the map has no room for.
    
    Fixes: 184c449f91fe ("net: Add support for XPS with QoS via traffic classes")
    Signed-off-by: Norbert Szetei <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
netfilter: ctnetlink: fix suspicious RCU usage in expect_iter_name [+ + +]
Author: Naman Gulati <[email protected]>
Date:   Sat Sep 12 01:10:51 2026 +0000

    netfilter: ctnetlink: fix suspicious RCU usage in expect_iter_name
    
    [ Upstream commit 207d591c353201f3bd3e0c89bb7d44a849c8fd59 ]
    
    expect_iter_name() is invoked by nf_ct_expect_iterate_net() under
    spin_lock_bh(&nf_conntrack_expect_lock). It does not hold
    rcu_read_lock().
    
    When accessing exp->helper with rcu_dereference() in syzbot's report,
    lockdep warns:
    
      =============================
      WARNING: suspicious RCU usage
      syzkaller #0 Not tainted
      -----------------------------
      net/netfilter/nf_conntrack_netlink.c:3393 suspicious rcu_dereference_check() usage!
    
      locks held by syz-executor381/5628: 2, last CPU#1:
       #0: ffffffff9aee42a0 (nfnl_subsys_ctnetlink_exp){+.+.}-{4:4},
           at: nfnetlink_rcv_msg+0xa69/0x12b0
       #1: ffffffff8ea74d58 (nf_conntrack_expect_lock){+...}-{3:3},
           at: nf_ct_expect_iterate_net+0x38/0x180
    
      Call Trace:
       <TASK>
       dump_stack_lvl+0xe8/0x150
       lockdep_rcu_suspicious+0x140/0x1d0
       expect_iter_name+0xfb/0x100
       nf_ct_expect_iterate_net+0xf2/0x180
       ctnetlink_del_expect+0x45d/0x640
       nfnetlink_rcv_msg+0xcc2/0x12b0
       netlink_rcv_skb+0x226/0x4a0
       nfnetlink_rcv+0x2b9/0x28c0
       netlink_unicast+0x7bd/0x940
       netlink_sendmsg+0x813/0xb40
       ____sys_sendmsg+0x54e/0x850
       ___sys_sendmsg+0x2a5/0x360
       __sys_sendmsg+0x2a5/0x360
       do_syscall_64+0x166/0x520
       entry_SYSCALL_64_after_hwframe+0x77/0x7f
    
    Use rcu_dereference_protected() with lockdep_is_held() on
    nf_conntrack_expect_lock instead, similar to expect_iter_me() in
    nf_conntrack_helper.c.
    
    Fixes: f01794106042 ("netfilter: nf_conntrack_expect: use expect->helper")
    Reported-by: [email protected]
    Closes: https://lore.kernel.org/netdev/[email protected]/T/#u
    Signed-off-by: Naman Gulati <[email protected]>
    Signed-off-by: Pablo Neira Ayuso <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

netfilter: flowtable: publish HW_DEAD after worker is done [+ + +]
Author: Jérémy Jean <[email protected]>
Date:   Tue Aug 18 20:00:15 2026 +0000

    netfilter: flowtable: publish HW_DEAD after worker is done
    
    [ Upstream commit d644b23afe1ef509c9961a6d84a093c2587edf02 ]
    
    flow_offload_work_del() sets NF_FLOW_HW_DEAD before the work handler
    clears NF_FLOW_HW_PENDING. Once a flow is both HW_DYING and HW_DEAD, a
    concurrent garbage collection pass can remove it and schedule it for RCU
    freeing.
    
    The offload worker holds neither an RCU read lock nor a reference to the
    flow. If it is preempted after publishing HW_DEAD, the RCU callback can
    free the flow before the worker resumes and clears HW_PENDING, resulting
    in a use-after-free.
    
    Move HW_DEAD publication to the common worker epilogue after the pending
    bit is cleared, making it the final flow access by destroy work. Order all
    preceding flow accesses before publishing the bit that allows garbage
    collection to free the object.
    
    Fixes: 2c8897953f3b ("netfilter: flowtable: Add pending bit for offload work")
    Assisted-by: Codex:gpt-5
    Signed-off-by: Jérémy Jean <[email protected]>
    Signed-off-by: Pablo Neira Ayuso <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

netfilter: ip6t_rpfilter: reject routes without inet6_dev [+ + +]
Author: Weiming Shi <[email protected]>
Date:   Sun Sep 6 16:44:10 2026 +0800

    netfilter: ip6t_rpfilter: reject routes without inet6_dev
    
    commit 1b9b5323725e458906c7620a3bc10398b51ad954 upstream.
    
    ip6_route_lookup() can return an error-free route whose rt6i_idev is
    NULL. Lowering an external nexthop device's MTU below IPV6_MIN_MTU tears
    down its inet6_dev while fib6_ifdown() leaves routes using nexthop objects
    in the FIB. An unprivileged user can construct this state with rtnetlink
    in a private user and network namespace, then trigger a NULL dereference
    through an IPv6 rpfilter lookup:
    
      Oops: general protection fault, probably for non-canonical address
      0xdffffc0000000000
      KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007]
      RIP: rpfilter_mt (net/ipv6/netfilter/ip6t_rpfilter.c:75)
      Call Trace:
      ip6t_do_table (net/ipv6/netfilter/ip6_tables.c:316)
      nf_hook_slow (net/netfilter/core.c:619)
      ipv6_rcv (net/ipv6/ip6_input.c:351)
      __netif_receive_skb_one_core (net/core/dev.c:6216)
      process_backlog (net/core/dev.c:6680)
      __napi_poll (net/core/dev.c:7739)
      net_rx_action (net/core/dev.c:7959)
      handle_softirqs (kernel/softirq.c:622)
      do_softirq.part.0 (kernel/softirq.c:523)
      __local_bh_enable_ip (kernel/softirq.c:450)
      __dev_queue_xmit (net/core/dev.c:4913)
      packet_sendmsg (net/packet/af_packet.c:3139)
      __sys_sendto (net/socket.c:2252)
      __x64_sys_sendto (net/socket.c:2259)
      do_syscall_64 (arch/x86/entry/syscall_64.c:94)
      entry_SYSCALL_64_after_hwframe (arch/x86/entry/entry_64.S:121)
      Kernel panic - not syncing: Fatal exception in interrupt
    
    Reject routes without an inet6_dev immediately after lookup. Such routes
    are not eligible for reverse-path filtering, and the check protects all
    later rt6i_idev dereferences.
    
    Fixes: e26f9a480fb6 ("netfilter: add ipv6 reverse path filter match")
    Reported-by: [email protected]
    Closes: https://lore.kernel.org/all/[email protected]/
    Suggested-by: Florian Westphal <[email protected]>
    Assisted-by: Claude:gpt-5
    Cc: [email protected]
    Signed-off-by: Weiming Shi <[email protected]>
    Signed-off-by: Pablo Neira Ayuso <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

netfilter: ip6t_rt: fix zero-address non-strict match out-of-bounds read [+ + +]
Author: Luxiao Xu <[email protected]>
Date:   Sun Sep 6 21:29:55 2026 +0800

    netfilter: ip6t_rt: fix zero-address non-strict match out-of-bounds read
    
    commit 82313c169eddc02b1bf5ba6b427803e272d3ec42 upstream.
    
    rt_mt6_check() permits rules to be configured with rtinfo->addrnr == 0
    even when address matching (IP6T_RT_FST_MASK) is requested.
    
    In the IP6T_RT_FST_NSTRICT path, rt_mt6() evaluates packet routing
    addresses against rtinfo->addrs[i] and terminates backwards at the bottom
    of the loop:
    
        if (ipv6_addr_equal(ap, &rtinfo->addrs[i])) {
            i++;
        }
        if (i == rtinfo->addrnr)
            break;
    
    When addrnr is 0, if the first packet address matches rtinfo->addrs[0],
    i is incremented to 1. Because i is now strictly greater than addrnr (0),
    the loop termination condition (i == rtinfo->addrnr) is bypassed and will
    never be satisfied.
    
    If a crafted IPv6 packet contains matching routing addresses, i will
    advance past IP6T_RT_HOPS (16). The subsequent call to ipv6_addr_equal()
    reads beyond struct ip6t_rt, triggering UBSAN/KASAN out-of-bounds warnings
    or kernel panics.
    
    Fix this by:
    1. Rejecting rules in rt_mt6_check() where IP6T_RT_FST_MASK is set but
       rtinfo->addrnr is zero.
    2. In rt_mt6(), moving the termination condition (i < rtinfo->addrnr)
       into the for-loop header condition and removing the backwards break
       at the end of the loop body.
    
    Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
    Cc: [email protected]
    Reported-by: Vega <[email protected]>
    Suggested-by: Florian Westphal <[email protected]>
    Assisted-by: LLM
    Signed-off-by: Luxiao Xu <[email protected]>
    Signed-off-by: Ren Wei <[email protected]>
    Signed-off-by: Pablo Neira Ayuso <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

netfilter: nf_tables: skip expired catchall elements on insert and delete [+ + +]
Author: Aohan Mei <[email protected]>
Date:   Mon Sep 14 19:51:47 2026 +0800

    netfilter: nf_tables: skip expired catchall elements on insert and delete
    
    commit 70194dc37670bd08e44b471389861cc01bd3a3c9 upstream.
    
    nft_setelem_catchall_insert() looks up duplicates with
    nft_set_elem_active() only, while nft_set_catchall_lookup() and the
    dump path additionally skip expired elements.
    
    Once a catchall element with a timeout expires, this predicate drift
    makes it invisible to userspace dumps, yet it still blocks
    re-insertion: with NLM_F_EXCL the request fails with -EEXIST, and
    without it the request reports success but silently inserts nothing.
    The stale entry only goes away when the (user-tunable) gc interval
    elapses, so the catchall rule may silently stop matching for an
    arbitrarily long time after its first expiration.
    
    The delete path shows the same drift: nft_setelem_catchall_deactivate()
    picks the first active-next entry in the catchall list, so with an
    expired entry still pending GC it retires the stale entry instead of
    the fresh one, and it deactivates an element that userspace no longer
    sees instead of failing with -ENOENT.
    
    Align both walks with the lookup and dump predicates: only an element
    that is active and not expired counts as a duplicate or delete
    candidate, using the per-netns timestamp taken at transaction start,
    in line with the set backend .insert/.deactivate and catchall GC sync
    paths.
    
    Reported-by: TencentOS Corvus AI <[email protected]>
    Cc: [email protected]
    Fixes: aaa31047a6d2 ("netfilter: nftables: add catch-all set element support")
    Assisted-by: CodeBuddy:Kimi-K3
    Signed-off-by: Aohan Mei <[email protected]>
    Signed-off-by: Pablo Neira Ayuso <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

netfilter: nfnetlink_queue: hold nfnl mutex in event notifier [+ + +]
Author: Florian Westphal <[email protected]>
Date:   Thu Sep 3 02:41:46 2026 +0200

    netfilter: nfnetlink_queue: hold nfnl mutex in event notifier
    
    [ Upstream commit 9461613afc59acef44a0071b0dd5075f6e993ffe ]
    
    We must serialize the release notifier and the config netlink function.
    A concurrent thread can issue close() which can call the release function
    while unrelated socket processes UNBIND request for same portid:
    
    Oops: general protection fault, [..]
    RIP: 0010:__instance_destroy+0x60/0x210 [nfnetlink_queue]
    Call Trace:
     nfqnl_recv_config+0x9b0/0xdc0 [nfnetlink_queue]
     nfnetlink_rcv_msg+0x7c2/0xeb0
     ? __pfx_nfnetlink_rcv_msg+0x10/0x10
    
    After this, parallel UNBIND and URELEASE events are impossible.
    
    This change isn't nice, but its the shortest fix given instances
    are not refcounted and the nfnetlink config callback drops the
    rcu read lock early due to need for sleeping allocations.
    
    Fixes: 7af4cc3fa158 ("[NETFILTER]: Add "nfnetlink_queue" netfilter queue handler over nfnetlink")
    Signed-off-by: Florian Westphal <[email protected]>
    Signed-off-by: Pablo Neira Ayuso <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

netfilter: nft_synproxy: use the family-aware checksum helper [+ + +]
Author: Karl Mehltretter <[email protected]>
Date:   Thu Sep 10 22:02:28 2026 +0200

    netfilter: nft_synproxy: use the family-aware checksum helper
    
    [ Upstream commit a311a898172743558b82f6035ef2aa8c310a4223 ]
    
    nft_synproxy_do_eval() verifies the TCP checksum before it switches on
    skb->protocol.  It uses nf_ip_checksum(), which constructs an IPv4
    pseudo header and relies on the IPv4 header checksum when folding the
    whole skb.  Neither operation is valid for an IPv6 packet.
    
    A correctly checksummed IPv6 segment can therefore fail verification
    when it reaches the hook as CHECKSUM_NONE or, at NF_INET_LOCAL_IN,
    CHECKSUM_COMPLETE.  nft_synproxy_do_eval() returns NF_DROP before
    nft_synproxy_eval_v6() can send a SYN-ACK.
    
    nft_synproxy_validate() deliberately admits NFPROTO_IPV6 and
    NFPROTO_INET, and the xtables counterpart ip6t_SYNPROXY.c already calls
    nf_ip6_checksum().
    
    Use nf_checksum() with nft_pf() so the checksum helper dispatches to the
    packet family's implementation.
    
    Fixes: ad49d86e07a4 ("netfilter: nf_tables: Add synproxy support")
    Assisted-by: LLM
    Signed-off-by: Karl Mehltretter <[email protected]>
    Signed-off-by: Pablo Neira Ayuso <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
netfs, afs: Fix symlink reading [+ + +]
Author: David Howells <[email protected]>
Date:   Tue Sep 15 17:21:59 2026 +0100

    netfs, afs: Fix symlink reading
    
    [ Upstream commit c51e89c5b5db3e1927738fad7cd18b66dde68aed ]
    
    Fix the reading of symlinks from the cache in afs by making netfslib trim
    the amount read down to i_size.  The problem is that afs sets the size of
    the iterator to the size of the buffer (PAGE_SIZE) so that the cache can
    round the read size up to the cache's DIO size.
    
    Note that this also impacts the reading of AFS mountpoints as they're just
    stored as symlinks with an odd file mode.
    
    Link: https://patch.msgid.link/[email protected]
    Fixes: c0410adf3da6 ("afs: Fix the locking used by afs_get_link()")
    Reviewed-by: Paulo Alcantara <[email protected]>
    cc: Paulo Alcantara <[email protected]>
    cc: Marc Dionne <[email protected]>
    cc: [email protected]
    cc: [email protected]
    cc: [email protected]
    Signed-off-by: Christian Brauner (Amutable) <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
netfs: Fix missing alloc tagging of direct mempool allocations [+ + +]
Author: Hao Ge <[email protected]>
Date:   Wed Sep 23 14:37:59 2026 +0800

    netfs: Fix missing alloc tagging of direct mempool allocations
    
    commit b78b728e21c32ec4c330b299f657fb1eb02dffc2 upstream.
    
    Commit 1d78d56c43ef ("netfs: Fix folio_queue ENOMEM in writeback by
    adding a mempool") added a mempool for the folio_queues and made the
    request, subrequest and folio_queue allocations distinguish between
    writeback and everything else.  Writeback is part of memory reclaim
    and must not fail due to ENOMEM, so it allocates under GFP_NOFS
    through mempool_alloc(), which may dip into the pool's reserve and,
    if that runs empty, wait for elements to be returned.  The
    GFP_KERNEL paths, which can return -ENOMEM to their callers, invoke
    the pool's ->alloc() callback directly instead.
    
    The direct call, however, skips the alloc_hooks() wrapper that the
    mempool_alloc() macro provides.  The pool callbacks, mempool_alloc_slab()
    and mempool_kmalloc(), call kmem_cache_alloc_noprof() and kmalloc_noprof()
    and rely on current->alloc_tag having been set by the caller.  With
    CONFIG_MEM_ALLOC_PROFILING_DEBUG=y this leads to
    
        current->alloc_tag not set
        WARNING: ./include/linux/alloc_tag.h:161 at __alloc_tagging_slab_alloc_hook
        alloc_tag was not set
        WARNING: ./include/linux/alloc_tag.h:166 at __alloc_tagging_slab_free_hook
    
    at allocation and free time respectively, as reported when reading
    files on a CIFS mount.  The allocations are also missing from
    /proc/allocinfo.
    
    Wrap the direct ->alloc() invocations in alloc_hooks() with a new
    mempool_alloc_noreserve() helper in include/linux/mempool.h, next to
    the other alloc_hooks()-wrapped macros such as mempool_alloc().  The
    GFP_KERNEL paths keep their failable allocation semantics, they just
    get tagged now.
    
    Fixes: 1d78d56c43ef ("netfs: Fix folio_queue ENOMEM in writeback by adding a mempool")
    Reported-by: Erhard Furtner <[email protected]>
    Closes: https://lore.kernel.org/all/[email protected]/
    Tested-by: Erhard Furtner <[email protected]>
    Suggested-by: Suren Baghdasaryan <[email protected]>
    Cc: [email protected]
    Signed-off-by: Hao Ge <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Acked-by: Vlastimil Babka (SUSE) <[email protected]>
    Signed-off-by: Christian Brauner (Amutable) <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

netfs: Fix netfs_read_gaps() to use separate sink folios [+ + +]
Author: David Howells <[email protected]>
Date:   Mon Sep 14 16:20:27 2026 +0100

    netfs: Fix netfs_read_gaps() to use separate sink folios
    
    [ Upstream commit fc3ae66514ca5e87251f79044236e3e8babd24a3 ]
    
    Fix netfs_read_gaps() to use separate folios rather than re-using a single
    sink folio to discard the unwanted data so that cifs checksum checking sees
    all the data that was fetched.
    
    Fixes: 7f84a7b9892d ("netfs: Make netfs_read_folio() handle streaming-write pages")
    Reported-by: Frank Sorenson <[email protected]>
    Closes: https://lore.kernel.org/r/[email protected]/
    Signed-off-by: David Howells <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Tested-by: Frank Sorenson <[email protected]>
    Reviewed-by: Paulo Alcantara <[email protected]>
    cc: Paulo Alcantara <[email protected]>
    cc: Namjae Jeon <[email protected]>
    cc: [email protected]
    cc: [email protected]
    cc: [email protected]
    Signed-off-by: Christian Brauner (Amutable) <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
nfc: fix use-after-free in nfc_get_local_general_bytes [+ + +]
Author: Luxiao Xu <[email protected]>
Date:   Wed Sep 9 13:19:24 2026 +0800

    nfc: fix use-after-free in nfc_get_local_general_bytes
    
    commit dcab71a7011918f6fdba7adcec02d217dcb84b8d upstream.
    
    Commit 6709d4b7bc2e ("net: nfc: Fix use-after-free caused by
    nfc_llcp_find_local") attempted to fix a use-after-free (UAF) issue by
    invoking nfc_llcp_local_put(local) after accessing local->gb. However,
    if the reference count drops to zero, local is freed immediately,
    leading to a use-after-free when callers access the returned pointer.
    Alternative approaches using dynamic allocation (e.g. kmemdup) introduced
    memory leaks because callers consistently treat the returned pointer as
    borrowed memory.
    
    Fix this properly by refactoring nfc_llcp_general_bytes() and
    nfc_get_local_general_bytes() to accept a caller-provided output buffer
    (out_gb) and its maximum length (gb_max_len). The general bytes are
    safely copied into out_gb before calling nfc_llcp_local_put(local),
    ensuring safe lifetime management without ownership transfer complications.
    
    Update all callers across drivers (microread, pn533, pn544, st21nfca,
    digital_dep, and nci) to provide their own destination buffers and pass
    them to nfc_get_local_general_bytes().
    
    Fixes: 6709d4b7bc2e ("net: nfc: Fix use-after-free caused by nfc_llcp_find_local")
    Cc: [email protected]
    Reported-by: Vega <[email protected]>
    Assisted-by: LLM
    Signed-off-by: Luxiao Xu <[email protected]>
    Signed-off-by: Ren Wei <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/3cbaac3bee23f8ff3a3284ed32d347696eb1d208.1788841683.git.rakukuip@gmail.com
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

nfc: llcp: drop truncated I/RR/RNR PDUs in nfc_llcp_recv_hdlc() [+ + +]
Author: Aamir Ahmed <[email protected]>
Date:   Tue Sep 15 19:54:27 2026 +0100

    nfc: llcp: drop truncated I/RR/RNR PDUs in nfc_llcp_recv_hdlc()
    
    commit 273f9d667cde649f8de9d72b1303cc2f4b658c50 upstream.
    
    nfc_llcp_recv_hdlc() reads the sequence byte skb->data[2], via
    nfc_llcp_ns()/nfc_llcp_nr(), before any length check. The receive path
    only guarantees the two-byte LLCP header -- __nfc_llcp_recv() checks it
    with pskb_may_pull() and nfc_llcp_recv_agf() admits two-byte inner PDUs
    -- so a two-byte I, RR or RNR PDU reads one byte of uninitialised skb
    tailroom. The byte becomes N(R)/N(S); a peer can already set those with
    a well-formed PDU, so this is acting on uninitialised memory, not new
    peer control.
    
    Guard the read with pskb_may_pull(), as commit 95674f506c63 ("nfc: llcp:
    reject PDUs shorter than the LLCP header") did for the two-byte header,
    so the sequence byte is present and linear before it is read. RR and RNR
    PDUs are LLCP_HEADER_SIZE + LLCP_SEQUENCE_SIZE bytes and an I PDU is
    longer, so no valid frame is rejected; a truncated PDU is malformed, so
    return without a DM reply.
    
    Fixes: d646960f7986 ("NFC: Initial LLCP support")
    Cc: [email protected]
    Assisted-by: LLM
    Signed-off-by: Aamir Ahmed <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/AS8P251MB0001789BBF04B72745C7D96BC8BA2@AS8P251MB0001.EURP251.PROD.OUTLOOK.COM
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

nfc: llcp: fix -ENOMEM on connect with zero-length service name [+ + +]
Author: Ömer Mete Kaya <[email protected]>
Date:   Wed Sep 9 15:16:22 2026 +0300

    nfc: llcp: fix -ENOMEM on connect with zero-length service name
    
    [ Upstream commit c04981e42d94f39c1dba965cc462a046e946a6c5 ]
    
    When service_name_len is 0, kmemdup() returns ZERO_SIZE_PTR which
    passes the NULL check, causing nfc_llcp_send_connect() to attempt
    building a zero-length service name TLV and fail with -ENOMEM.
    
    Fix by setting service_name to NULL directly when service_name_len is 0.
    
    Fixes: d646960f7986 ("NFC: Initial LLCP support")
    Signed-off-by: Ömer Mete Kaya <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

nfc: llcp: Fix list corruption / refcount desync in nfc_llcp_recv_dm() [+ + +]
Author: Aldo Ariel Panzardo <[email protected]>
Date:   Thu Jul 16 20:26:57 2026 -0300

    nfc: llcp: Fix list corruption / refcount desync in nfc_llcp_recv_dm()
    
    [ Upstream commit bf1460acdf8cf5a07c819f59785d40f20d113099 ]
    
    nfc_llcp_recv_dm() handles DM(NOBOUND)/DM(REJ) for a socket that is still
    linked on local->connecting_sockets: it looks the socket up with
    nfc_llcp_connecting_sock_get(), sets sk->sk_state = LLCP_CLOSED and
    returns, without taking the socket lock and without unlinking the socket
    from the connecting_sockets list.
    
    llcp_sock_release() selects the list to unlink from by sk_state: a socket
    in LLCP_CONNECTING is unlinked from connecting_sockets, otherwise from the
    sockets list.  Because recv_dm left the socket physically on
    connecting_sockets but in the LLCP_CLOSED state, release() takes the else
    branch and calls nfc_llcp_sock_unlink(&local->sockets, sk).  That runs
    sk_del_node_init() while holding sockets.lock, i.e. it removes the socket
    from the connecting_sockets hlist under the wrong lock.  A concurrent
    connect() linking another socket onto connecting_sockets under
    connecting_sockets.lock then mutates the same hlist unserialized, which
    corrupts the list and desyncs the sk_add_node()/sk_del_node_init()
    sock_hold()/__sock_put() pairing.  An unprivileged local process holding
    LLCP sockets, with the DM supplied by the remote peer over an established
    LLCP link, can drive this to leak kernel sockets without bound (the
    mis-decrement goes through the non-freeing __sock_put() path, so the
    object is never released), leading to memory exhaustion / DoS.
    
    This is the same class of bug that was fixed in the sibling handler
    nfc_llcp_recv_cc() by commit b493ea2765cc ("nfc: llcp: Fix use-after-free
    race in nfc_llcp_recv_cc()"); recv_dm did not receive the equivalent fix.
    
    Fix it the same way: take lock_sock(), re-check that the socket is still
    hashed (release() may have won the race), and for the NOBOUND/REJ case
    unlink it from connecting_sockets before moving it to LLCP_CLOSED.  The
    unlink drops the connecting_sockets membership reference via
    sk_del_node_init(), leaving the socket unhashed, so the later
    nfc_llcp_sock_unlink() in llcp_sock_release() becomes a no-op and no
    double put occurs.
    
    Fixes: a69f32af86e3 ("NFC: Socket linked list")
    Signed-off-by: Aldo Ariel Panzardo <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

nfc: llcp: Fix race condition in accept_queue lifecycle [+ + +]
Author: Lee Jones <[email protected]>
Date:   Wed Sep 2 12:30:31 2026 +0000

    nfc: llcp: Fix race condition in accept_queue lifecycle
    
    [ Upstream commit c3eef2f988a3db9690369d7cef9a3344dd9788d3 ]
    
    In nfc_llcp_socket_release(), sockets and listener accept queues are
    walked under the local sockets rwlock and bh_lock_sock().  However,
    bh_lock_sock() does not synchronise against process-context lock_sock()
    held by nfc_llcp_accept_dequeue() during accept().  Because
    socket_release() does not check sock_owned_by_user(), both paths can
    concurrently unlink and release the same child socket, resulting in
    use-after-free or a NULL pointer dereference of child->parent in
    nfc_llcp_accept_unlink().
    
    Fix this synchronisation race by having nfc_llcp_socket_release() use
    process-context lock_sock() instead of bh_lock_sock():
    
    1. Pop sockets from the local sockets list under the write lock using
       nfc_llcp_sock_list_pop() so lock_sock() can be acquired without
       holding the rwlock.
    
    2. Because lock_sock() can sleep, defer the final release of the
       nfc_llcp_local structure to a workqueue (release_work).  This avoids
       a sleeping-in-atomic bug when the last local reference is dropped
       from softirq context.  Additionally, hold a single device reference
       on local from registration until final destruction.
    
    3. In nfc_llcp_local_get(), use kref_get_unless_zero() to prevent
       resurrecting a local object whose teardown has been scheduled.
    
    4. In llcp_sock_accept(), verify that the listener socket state is still
       LLCP_LISTEN after waking from schedule_timeout() to prevent hangs if
       the listener is closed concurrently.
    
    5. When unlinking unaccepted child sockets during listener release,
       unlink them from local->sockets, call sock_orphan(), and drop their
       initial sk_alloc creation reference via sock_put().
    
    6. Make nfc_llcp_accept_unlink() idempotent by guarding parent access with
       a NULL check.
    
    Fixes: 50b78b2a6500 ("NFC: Fix sleeping in atomic when releasing socket")
    Signed-off-by: Lee Jones <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

nfc: llcp: fix sdreq TLV list leak on parse/alloc/send failure [+ + +]
Author: Cong Nguyen <[email protected]>
Date:   Mon Sep 14 19:11:29 2026 +0700

    nfc: llcp: fix sdreq TLV list leak on parse/alloc/send failure
    
    [ Upstream commit 66f4300206b82b0b143ef0d9be90cd8d29f23a47 ]
    
    nfc_genl_llc_sdreq() builds a list of TLV nodes while walking nested
    netlink attrs, but 3 error paths (nested-attr parse failure, TLV alloc
    ENOMEM, nfc_llcp_send_snl_sdreq() failure) all skip freeing what was
    already queued.
    
    Route them through a new free_list label, mirroring the SDRES path in
    the same file which already does this. Harmless on the success path
    too -- send_snl_sdreq() drains the list as it moves nodes, so it's
    already empty by the time free_list runs.
    
    Fixes: d9b8d8e19b07 ("NFC: llcp: Service Name Lookup netlink interface")
    Assisted-by: Claude:claude-opus-4
    Signed-off-by: Cong Nguyen <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

nfc: llcp: fix slab-out-of-bounds reads when logging service names [+ + +]
Author: Ömer Mete Kaya <[email protected]>
Date:   Tue Sep 8 19:18:01 2026 +0300

    nfc: llcp: fix slab-out-of-bounds reads when logging service names
    
    [ Upstream commit 7dcf371a35632f035baf77bcf2c129165f772ce4 ]
    
    nfc_llcp_wks_sap() and nfc_llcp_build_sdreq_tlv() pass non-null-
    terminated strings to pr_debug() using the %s format specifier.
    The buffers are allocated via kmemdup() or come from netlink
    attributes and are not guaranteed to be null-terminated, causing
    __dynamic_pr_debug() to read beyond the allocated region:
    
      KASAN: slab-out-of-bounds Read in __dynamic_pr_debug
    
    Fix both call sites by using %.*s with the explicit length to limit
    the output to the actual length of the string.
    
    Fixes: d9b8d8e19b07 ("NFC: llcp: Service Name Lookup netlink interface")
    Reported-by: [email protected]
    Closes: https://syzkaller.appspot.com/bug?extid=1e3df0852e82c21ca418
    Signed-off-by: Ömer Mete Kaya <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

nfc: llcp: fix WKS SAP hijacking via prefix match in nfc_llcp_wks_sap() [+ + +]
Author: Ömer Mete Kaya <[email protected]>
Date:   Wed Sep 9 15:14:31 2026 +0300

    nfc: llcp: fix WKS SAP hijacking via prefix match in nfc_llcp_wks_sap()
    
    [ Upstream commit 408cff6bd60636df201274d320edfdfde9ed41db ]
    
    nfc_llcp_wks_sap() compares only service_name_len bytes, so a short
    service_name like "u" matches longer WKS strings like "urn:nfc:sn:snep".
    Fix by requiring exact length match before strncmp().
    
    Fixes: d646960f7986 ("NFC: Initial LLCP support")
    Signed-off-by: Ömer Mete Kaya <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

nfc: nfcmrvl: validate helper command length before pull [+ + +]
Author: Pengpeng Hou <[email protected]>
Date:   Wed Jul 15 16:43:25 2026 +0800

    nfc: nfcmrvl: validate helper command length before pull
    
    [ Upstream commit 686f942332b1667f13f3b8d6a2f50bcfbf42e277 ]
    
    The firmware download receive path removes the NCI data header and
    reads the helper command before validating the remaining packet length.
    A short frame can therefore reach the data access before the malformed
    packet is rejected.
    
    Validate the complete helper command length before stripping the NCI
    data header.
    
    Fixes: 3194c6870158 ("NFC: nfcmrvl: add firmware download support")
    Signed-off-by: Pengpeng Hou <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

nfc: pn533: fix OOB read in pn533_acr122_is_rx_frame_valid() [+ + +]
Author: Deepanshu Kartikey <[email protected]>
Date:   Wed Sep 23 09:26:27 2026 +0530

    nfc: pn533: fix OOB read in pn533_acr122_is_rx_frame_valid()
    
    [ Upstream commit b61732f47316d45f27706db7812950145d3327b5 ]
    
    frame->ccid.datalen is read directly from the USB response frame
    and used, unchecked, as an index into frame->data[]. A malicious or
    malfunctioning device can set this field to an arbitrary value,
    causing the driver to read far outside the received buffer.
    
    Bound ccid.datalen against the maximum possible ACR122 frame size
    before using it. This replaces the existing datalen == 0 check,
    since datalen < 2 already covers that case and additionally
    rejects datalen == 1, which would still underflow the
    "datalen - 2" offset used below.
    
    Fixes: 9815c7cf22da ("NFC: pn533: Separate physical layer from the core implementation")
    Reported-by: [email protected]
    Closes: https://syzkaller.appspot.com/bug?extid=1853daab1a47603d4678
    Tested-by: [email protected]
    Assisted-by: LLM
    Signed-off-by: Deepanshu Kartikey <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

nfc: port100: reject frames whose declared length exceeds the received data [+ + +]
Author: Doruk Tan Ozturk <[email protected]>
Date:   Sat Jul 11 14:36:51 2026 +0200

    nfc: port100: reject frames whose declared length exceeds the received data
    
    commit 092c6a605cbd6414ef499834c2e0da69c2c3388e upstream.
    
    port100_recv_response() passes the URB transfer buffer to
    port100_rx_frame_is_valid(), which checksums le16_to_cpu(frame->datalen)
    bytes of frame->data. datalen is a 16-bit field supplied by the device
    and is never checked against the number of bytes actually received
    (urb->actual_length), so a device reporting a datalen larger than the
    received frame makes port100_data_checksum() read out of bounds past the
    transfer buffer.
    
    Reject a response whose declared frame size does not fit the received
    length before validating it.
    
    Found by 0sec (https://0sec.ai) using automated source analysis; the
    missing bound is evident from source. Compile-tested.
    
    Fixes: 562d4d59b8a1 ("NFC: Sony Port-100 Series driver")
    Cc: [email protected]
    Assisted-by: 0sec:claude-opus-4-8
    Signed-off-by: Doruk Tan Ozturk <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

nfc: st21nfca: validate ISO15693 inventory length [+ + +]
Author: Pengpeng Hou <[email protected]>
Date:   Sun Aug 30 21:29:58 2026 +0800

    nfc: st21nfca: validate ISO15693 inventory length
    
    [ Upstream commit 7f2ea5ed588c03d481f0301e6c3d4240132383fb ]
    
    The ISO15693 inventory helper removes a two-byte prefix without checking
    that it exists, then accepts a one-byte remainder before reading data[1] as
    the DSFID.
    
    Require the prefix and at least two remaining bytes before copying the UID
    data and reading the DSFID.
    
    Fixes: 7974728094d3 ("NFC: st21nfca: Add ISO15693 Reader/Writer support")
    Signed-off-by: Pengpeng Hou <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

nfc: st21nfca: validate received frame size [+ + +]
Author: Pengpeng Hou <[email protected]>
Date:   Wed Jul 15 16:44:05 2026 +0800

    nfc: st21nfca: validate received frame size
    
    [ Upstream commit a653c01ce447f10c36b901646888c0330363af4f ]
    
    st21nfca_hci_i2c_repack() trims a received frame at its EOF marker
    before removing byte stuffing.  It then assumes the truncated frame
    contains the LLC header and two CRC bytes, and it unconditionally reads
    the byte after an escape marker.
    
    A malformed frame can place EOF immediately after the start marker or can
    end its data portion with an escape marker.  The former leaves too few
    bytes for check_crc(), while the latter makes the unstuffing loop read past
    the current skb length.
    
    Require the minimum framing bytes both before and after unstuffing.  Use
    separate input and output cursors while removing byte stuffing, and reject
    an escape marker without its encoded byte.  This keeps malformed frames
    within the received frame boundary before CRC processing.
    
    Fixes: 3096e25a3e40 ("NFC: st21nfca: Fix incorrect byte stuffing revocation")
    Signed-off-by: Pengpeng Hou <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

nfc: trf7970a: power down on startup RX gain failure [+ + +]
Author: Myeonghun Pak <[email protected]>
Date:   Sun Sep 13 00:26:25 2026 -0400

    nfc: trf7970a: power down on startup RX gain failure
    
    commit d2acbde7e67df44efa8f0963462d1192e7694ffc upstream.
    
    trf7970a_startup() powers up the device before applying the optional RX
    gain reduction. If the register read or write fails, it returns without
    undoing that power-up. Probe's unwind only drops the separate regulator
    references acquired by probe, leaving the additional VIN enable from
    startup unbalanced. The system resume caller also has no power-down on
    this error.
    
    Call trf7970a_power_down() before returning the RX gain error to deassert
    the enable GPIOs, release the startup VIN reference and restore the
    powered-off state. Runtime PM has not been enabled yet, so the full
    shutdown helper is not appropriate here. Preserve the original SPI error.
    
    This issue was identified during our ongoing static-analysis research while
    reviewing kernel code.
    
    Fixes: 5d69351820ea ("NFC: trf7970a: Create device-tree parameter for RX gain reduction")
    Cc: [email protected]
    Assisted-by: OpenAI:GPT-5.6
    Co-developed-by: Ijae Kim <[email protected]>
    Signed-off-by: Ijae Kim <[email protected]>
    Signed-off-by: Myeonghun Pak <[email protected]>
    Reviewed-by: Paul Geurts <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

nfc: virtual_ncidev: Add missing ioctl compat handler [+ + +]
Author: Chris Gellermann <[email protected]>
Date:   Fri Sep 4 18:42:52 2026 +0200

    nfc: virtual_ncidev: Add missing ioctl compat handler
    
    [ Upstream commit 51814683e28fc64eceb415962376956c3cfc75a7 ]
    
    The compat handler for ioctls to the virtual nci device is missing. So,
    nci-specific ioctls of a compat task return with -1 and errno set to
    ENOTTY. Add a handler.
    
    The handling of an ioctl() call of a compat task to get the index of
    virtual nci device (IOCTL_GET_NCIDEV_IDX) lands in the default case of
    the ioctl compat handler (see fs/ioctl.c):
    
    COMPAT_SYSCALL_DEFINE3(ioctl, ...)
    {
            ...
            default:
                    error = do_vfs_ioctl(fd_file(f), fd, cmd, ...);
                    if (error != -ENOIOCTLCMD)
                            break;
    
                    if (fd_file(f)->f_op->compat_ioctl)
                            error = fd_file(f)->f_op->compat_ioctl(fd_file(f), cmd, arg);
                    if (error == -ENOIOCTLCMD)
                            error = -ENOTTY;
            ...
    }
    
    There, do_vfs_ioctl() returns -ENOIOCTLCMD and compat_ioctl is not
    set for virtual_ncidev_fops, i.e. f_op->compat_ioctl == NULL. So, the
    ioctl() syscall returns with -1 and errno set to ENOTTY to the compat
    task.
    
    To fix this, use the compat_ptr_ioctl helper for compat handling here.
    It shall be used for ioctls that "either ignore the argument or pass a
    pointer to a compatible data type". The driver's sole ioctl takes a user
    void pointer and copies nfc_dev->idx to it, a 4-byte integer across all
    ABIs.
    
    This issue has been found by running the nci_dev kernel selftest as
    rv64 binary on top of a CHERI kernel, where the ioctl() ends up in
    the ioctl compat handler, similar to a 32-bit application on top of a
    64-bit kernel.
    
    Fixes: e624e6c3e777 ("nfc: Add a virtual nci device driver")
    Signed-off-by: Chris Gellermann <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
nfp: hold IPsec RX state under the XArray lock [+ + +]
Author: Sang-Hoon Choi <[email protected]>
Date:   Tue Sep 22 03:15:59 2026 +0900

    nfp: hold IPsec RX state under the XArray lock
    
    [ Upstream commit 1a983a4e14c635c40354be110cd9a1a5c94e01e6 ]
    
    nfp_net_ipsec_rx() drops the XArray lock before taking a reference to the
    xfrm_state it found. The delete path can erase the entry and drop the last
    state reference in that interval. RX can then try to increment a zero
    refcount after the state has been queued for destruction.
    
    The driver queues firmware invalidation asynchronously; the delete path
    does not wait for the command to complete or drain pending RX processing.
    
    The XFRM garbage collector waits for an RCU grace period before freeing
    the state. That delays reclamation but does not make acquiring a reference
    from zero valid.
    
    Take the xfrm_state reference before releasing the XArray lock so
    xa_erase() cannot run between lookup and reference acquisition.
    
    Fixes: 57f273adbcd4 ("nfp: add framework to support ipsec offloading")
    Reported-by: Changyul Lee <[email protected]>
    Signed-off-by: Sang-Hoon Choi <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
ocfs2: make ocfs2_calc_xattr_init() return void [+ + +]
Author: Joseph Qi <[email protected]>
Date:   Fri Sep 4 10:37:51 2026 +0800

    ocfs2: make ocfs2_calc_xattr_init() return void
    
    [ Upstream commit 525c0edc032b3297d0c1056cf1fa20cf1f9e6184 ]
    
    ocfs2_calc_xattr_init() used to read the default ACL off the parent inode
    itself, so it could return an error from ocfs2_xattr_get_nolock().  Commit
    bd7c05fb4a47 ("ocfs2: fix circular locking dependency in
    ocfs2_init_acl()") moved that lookup before the transaction starts and
    deleted the error path, but left the now vestigial 'int ret = 0'
    declaration and both 'return ret' statements behind, along with an
    unreachable error branch in ocfs2_mknod().
    
    Drop the leftover variable and convert the return type to void, so the
    callee states that it always succeeds and the caller no longer carries a
    check that can never trigger.
    
    No functional change.
    
    Link: https://lore.kernel.org/[email protected]
    Fixes: bd7c05fb4a47 ("ocfs2: fix circular locking dependency in ocfs2_init_acl()")
    Signed-off-by: Joseph Qi <[email protected]>
    Signed-off-by: Andrew Morton <[email protected]>
    Reported-by: kernel test robot <[email protected]>
    Closes: https://lore.kernel.org/oe-kbuild-all/[email protected]/
    Cc: Mark Fasheh <[email protected]>
    Cc: Joel Becker <[email protected]>
    Cc: Junxiao Bi <[email protected]>
    Cc: Changwei Ge <[email protected]>
    Cc: Jun Piao <[email protected]>
    Cc: Heming Zhao <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
octeontx2-af: Fix memory scaling limitation in SR-IOV mode [+ + +]
Author: Ratheesh Kannoth <[email protected]>
Date:   Wed Sep 16 07:51:11 2026 +0530

    octeontx2-af: Fix memory scaling limitation in SR-IOV mode
    
    [ Upstream commit d06f2ebf67ff2962fe00d687e4f0d4703eb41a12 ]
    
    The original code used DMA_ATTR_FORCE_CONTIGUOUS, which could exhaust
    the CMA pool when a large number of VFs were requested.
    
    Fix this by switching to the DMA streaming API. This is equivalent on
    Octeon platforms, which provide full I/O coherency via the SMMU.
    
    Cc: Leon Romanovsky <[email protected]>
    Fixes: 73d33dbc0723 ("octeontx2-af: Use DMA_ATTR_FORCE_CONTIGUOUS attribute in DMA alloc")
    Signed-off-by: Ratheesh Kannoth <[email protected]>
    Reviewed-by: Leon Romanovsky <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

octeontx2-af: use seq_file for rsrc_alloc debugfs [+ + +]
Author: Heyang Tan <[email protected]>
Date:   Mon Sep 14 10:05:21 2026 +0800

    octeontx2-af: use seq_file for rsrc_alloc debugfs
    
    [ Upstream commit 39c6580765dad6477fb2637f6f616e0d276aae65 ]
    
    The rsrc_alloc debugfs reader writes rows directly to userspace without
    respecting the caller's read count. It also uses the current row length as
    the userspace stride, which can corrupt output when rows have different
    widths.
    
    Use seq_file to handle userspace buffer sizes, offsets, and partial reads,
    and write output columns directly to the seq_file buffer.
    
    Fixes: 23205e6d06d4 ("octeontx2-af: Dump current resource provisioning status")
    Signed-off-by: Heyang Tan <[email protected]>
    Reviewed-by: Ratheesh Kannoth <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
ovl: fix UAF in ovl_do_mkdir() debug print [+ + +]
Author: Amir Goldstein <[email protected]>
Date:   Mon Sep 21 12:40:13 2026 +0200

    ovl: fix UAF in ovl_do_mkdir() debug print
    
    [ Upstream commit ae146bc1abdeb4607abf2975b858c053024e8ac1 ]
    
    ovl_do_mkdir() prints the input dentry with %pd after vfs_mkdir().
    Since commit fe497f0759e0 ("VFS: change vfs_mkdir() to unlock on
    failure."), vfs_mkdir() calls end_creating() on the input dentry on
    failure and may replace it on success, so the post-call %pd can
    use-after-free the dentry when CONFIG_OVERLAY_FS_DEBUG is enabled.
    
    Print the dentry before the call and only the result afterward.
    
    Reported-by: [email protected]
    Closes: https://syzkaller.appspot.com/bug?extid=ced26b784bf977d223dd
    Fixes: fe497f0759e0 ("VFS: change vfs_mkdir() to unlock on failure.")
    Signed-off-by: Amir Goldstein <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Christian Brauner (Amutable) <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
ovpn: always unhash old VPN addresses before rehashing [+ + +]
Author: Ralf Lici <[email protected]>
Date:   Fri Aug 28 15:00:06 2026 +0200

    ovpn: always unhash old VPN addresses before rehashing
    
    [ Upstream commit b43beccb3713fafada57814b0a652f4a876eb75f ]
    
    ovpn_peer_hash_vpn_ip updates the per-peer VPN address hash entries
    after userspace changes a peer VPN address. The current code removes an
    old hash entry only when the new address for that family is not the
    unspecified address.
    
    When an address is cleared to 0.0.0.0 or ::, its hash node therefore
    remains linked in the bucket selected by the old address. The address
    comparison performed during lookup prevents the old address from
    matching, but the table retains a stale entry until the peer is removed
    or another address is configured for that family.
    
    Always remove both old VPN address hash entries before conditionally
    adding the currently configured addresses back. This ensures that a
    cleared address leaves its hash node unhashed.
    
    Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink")
    Signed-off-by: Ralf Lici <[email protected]>
    Signed-off-by: Antonio Quartulli <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

ovpn: preserve IPv6 scope id for netlink peer endpoints [+ + +]
Author: Ralf Lici <[email protected]>
Date:   Fri Aug 28 16:50:22 2026 +0200

    ovpn: preserve IPv6 scope id for netlink peer endpoints
    
    [ Upstream commit 7a6d08ee0f0e30023d18779bb314db8fd9a3b6d4 ]
    
    ovpn accepts OVPN_A_PEER_REMOTE_IPV6_SCOPE_ID and reports
    bind->remote.in6.sin6_scope_id in peer dumps, but the netlink endpoint
    parser never copied the attribute into the sockaddr_in6 used to create or
    update the peer bind.
    
    As a result, an IPv6 link-local remote endpoint configured through
    netlink loses its interface scope, unlike on the peer float path where
    ipv6_iface_scope_id populates the field. The UDPv6 output path then
    builds a flow with flowi6_oif set to zero and route lookup can fail or
    select the wrong interface.
    
    Copy the scope id when parsing non-v4-mapped IPv6 remote endpoints. The
    existing precheck already rejects the scope-id attribute for IPv4 and
    v4-mapped IPv6 remotes.
    
    Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink")
    Signed-off-by: Ralf Lici <[email protected]>
    Signed-off-by: Antonio Quartulli <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

ovpn: reject duplicate peer VPN addresses [+ + +]
Author: Ralf Lici <[email protected]>
Date:   Fri Aug 28 15:00:07 2026 +0200

    ovpn: reject duplicate peer VPN addresses
    
    [ Upstream commit d25e885b31a0f2808d936f95c9a558a8a792669b ]
    
    In MP mode, ovpn uses the peer VPN addresses as lookup keys for
    selecting the peer that should receive an outgoing tunnel packet.
    However, the netlink peer configuration path does not currently reject
    duplicate VPN addresses.
    
    If two peers are configured with the same VPN address, both can be
    inserted in the VPN address hash table and lookups return whichever peer
    is found first. This makes peer selection ambiguous and dependent on
    hash insertion order.
    
    Reject peer creation or update when the resulting VPN address is already
    assigned to another peer. Ignore unspecified addresses because those are
    not inserted in the VPN address hash tables.
    
    This changes such configurations from being accepted to being rejected,
    but they have never worked reliably because peer selection is ambiguous.
    
    Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink")
    Signed-off-by: Ralf Lici <[email protected]>
    Signed-off-by: Antonio Quartulli <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

ovpn: reject invalid peer VPN addresses [+ + +]
Author: Ralf Lici <[email protected]>
Date:   Fri Aug 28 15:00:09 2026 +0200

    ovpn: reject invalid peer VPN addresses
    
    [ Upstream commit 5940f3407b78062442cb01f541ef6eed709fc380 ]
    
    In MP mode, ovpn uses peer VPN addresses as lookup keys for selecting
    the peer that should receive outgoing tunnel packets. The netlink
    configuration path currently accepts address values that cannot sensibly
    identify a VPN peer, such as multicast, broadcast or loopback addresses.
    
    Reject invalid peer VPN addresses when creating or updating an MP peer.
    Keep accepting the unspecified address as the internal unset value,
    provided that at least one VPN address family remains configured.
    
    Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink")
    Signed-off-by: Ralf Lici <[email protected]>
    Signed-off-by: Antonio Quartulli <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

ovpn: reject multipeer peers without VPN addresses [+ + +]
Author: Ralf Lici <[email protected]>
Date:   Fri Aug 28 15:00:08 2026 +0200

    ovpn: reject multipeer peers without VPN addresses
    
    [ Upstream commit 025af3a0a892514f9f27f186338ba3d44365547a ]
    
    In MP mode, ovpn uses the peer VPN addresses to select the peer for
    outgoing tunnel packets. Peer creation currently requires a VPN IPv4 or
    IPv6 attribute, but it only checks for the presence of the attribute and
    not for a usable address value.
    
    This allows userspace to create an MP peer with only unspecified VPN
    addresses, or to update an existing peer so that both VPN address
    families become unspecified. Such a peer cannot be selected through the
    VPN address hash tables.
    
    Reject MP peer creation or update when the resulting peer would not have
    at least one VPN address configured.
    
    This changes such configurations from being accepted to being rejected,
    but they have never been usable because the peer cannot be selected
    through the VPN address hash tables.
    
    Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink")
    Signed-off-by: Ralf Lici <[email protected]>
    Signed-off-by: Antonio Quartulli <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

ovpn: replace bind when clearing stale local source [+ + +]
Author: Ralf Lici <[email protected]>
Date:   Fri Aug 28 16:50:27 2026 +0200

    ovpn: replace bind when clearing stale local source
    
    [ Upstream commit 7d8104988f423572df1f3347ce578037b1043f34 ]
    
    The UDP output fallback clears bind->local in place when the remembered
    source address is no longer usable. The bind is RCU-published and read
    locklessly by concurrent TX, so an IPv6 reader can observe a torn
    address.
    
    Retry the route lookup with source address autoselection without
    modifying the bind. After a successful lookup, revalidate the bind and
    route key under peer->lock, reset the dst cache, and best-effort publish
    a replacement bind with a wildcard local address.
    
    Do not cache the resolved dst when clearing the local source. Replacing
    the source invalidates all per-CPU cache entries, while
    dst_cache_set_ip4 and dst_cache_set_ip6 update only the current CPU
    slot. The current packet can still use the resolved route; if bind
    allocation fails, a later cache miss retries the repair.
    
    Fixes: 08857b5ec5d9 ("ovpn: implement basic TX path (UDP)")
    Signed-off-by: Ralf Lici <[email protected]>
    Signed-off-by: Antonio Quartulli <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

ovpn: replace bind when learning local endpoint [+ + +]
Author: Ralf Lici <[email protected]>
Date:   Fri Aug 28 16:50:26 2026 +0200

    ovpn: replace bind when learning local endpoint
    
    [ Upstream commit aea934a221ec6a867221e5b765f65f1857befd53 ]
    
    struct ovpn_bind is published through peer->bind with RCU, but local
    endpoint learning updates bind->local in place under peer->lock. UDP TX
    reads the field without that lock. In particular, a concurrent IPv6
    update can therefore result in a torn address read.
    
    Use ovpn_peer_reset_sockaddr to publish a replacement bind when learning
    a new local endpoint, just as a remote endpoint change does. Preserve
    the current remote address and reset the dst cache only after the new
    bind has been published successfully.
    
    Track remote endpoint changes separately so that float notification and
    transport-address rehashing remain limited to actual peer floats.
    
    Fixes: f0281c1d3732 ("ovpn: add support for updating local or remote UDP endpoint")
    Signed-off-by: Ralf Lici <[email protected]>
    Signed-off-by: Antonio Quartulli <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

ovpn: skip UDP source validation for unspecified addresses [+ + +]
Author: Ralf Lici <[email protected]>
Date:   Fri Aug 28 16:50:23 2026 +0200

    ovpn: skip UDP source validation for unspecified addresses
    
    [ Upstream commit 77393b4d72dfeb764b2af2b848acc659f6fcfd0a ]
    
    ovpn validates the cached local UDP source address before reusing or
    refreshing a peer dst cache. This is only meaningful when a concrete
    source address is selected.
    
    For IPv6, calling ipv6_chk_addr with :: checks whether the unspecified
    address itself is configured on the host. A peer may legitimately have
    bind->local.ipv6 set to :: when no local endpoint was configured or
    after a stale learned address was cleared. In that case the source
    should be left unspecified and selected by ip6_dst_lookup_flow().
    
    For IPv4, inet_confirm_addr(..., local = 0, ...) asks for local address
    autoselection rather than validating a chosen source. Skip the precheck
    there as well and let ip_route_output_flow select or reject the source.
    
    Only validate non-zero/non-any source addresses.
    
    Fixes: 08857b5ec5d9 ("ovpn: implement basic TX path (UDP)")
    Signed-off-by: Ralf Lici <[email protected]>
    Signed-off-by: Antonio Quartulli <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

ovpn: track UDP socket route key for peer dst cache [+ + +]
Author: Ralf Lici <[email protected]>
Date:   Fri Aug 28 16:50:24 2026 +0200

    ovpn: track UDP socket route key for peer dst cache
    
    [ Upstream commit 7c66b7a4ae80a9309e6dc1d24b7b6b897e6348eb ]
    
    ovpn stores the route used to transmit UDP packets in a per-peer dst
    cache. A cached dst is only valid for the route lookup inputs used when
    it was resolved.
    
    Some of those inputs are mutable while userspace still owns the UDP
    socket. In particular, changes to the socket mark or UDP source port do
    not invalidate ovpn's peer dst cache, so ovpn can keep using a route
    selected with an old socket route key.
    
    Replace the cached mark with a route key containing the socket-owned
    lookup inputs currently used by ovpn, and reset the peer dst cache when
    the key changes. Before storing a newly looked-up dst, recheck the route
    key under the peer lock so a dst resolved for stale socket state is not
    published.
    
    Fixes: 08857b5ec5d9 ("ovpn: implement basic TX path (UDP)")
    Signed-off-by: Ralf Lici <[email protected]>
    Signed-off-by: Antonio Quartulli <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

ovpn: validate peer state before caching UDP dst [+ + +]
Author: Ralf Lici <[email protected]>
Date:   Fri Aug 28 16:50:25 2026 +0200

    ovpn: validate peer state before caching UDP dst
    
    [ Upstream commit fa603710bdb9aea33c0d9cc2c05ed24d84f58753 ]
    
    UDP route lookup runs without peer->lock while the bind is protected by
    RCU. The route key is snapshotted separately. Either can change while
    the lookup is in progress.
    
    The TX path currently checks only the route key before publishing the
    looked-up dst. If the bind changes but the route key does not, a dst
    resolved from the old endpoint can be installed in the cache after the
    bind replacement.
    
    Compare both the bind pointer and the route key under peer->lock before
    updating the cache. The RCU read-side critical section keeps the old
    bind alive throughout the lookup, so pointer identity is sufficient to
    detect a replacement.
    
    Fixes: f0281c1d3732 ("ovpn: add support for updating local or remote UDP endpoint")
    Signed-off-by: Ralf Lici <[email protected]>
    Signed-off-by: Antonio Quartulli <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
packet: use ubuf_info completion for TX_RING packets [+ + +]
Author: Willem de Bruijn <[email protected]>
Date:   Fri Sep 18 20:47:29 2026 -0400

    packet: use ubuf_info completion for TX_RING packets
    
    commit 9518405613863d0bf0700927a21367f5942cf058 upstream.
    
    tpacket_snd sends skbs with frags pointing into its ring slots. Slots
    are released when skb->destructor is called.
    
    A call to skb_orphan calls skb->destructor before the skb is freed.
    This can cause the slot to be reused while still linked into the skb.
    
    Switch to standard zerocopy completion (ubuf_info) so the slot is only
    released once all references to the payload are freed or copied.
    Restore skb->destructor to standard sock_wfree.
    
    The ubuf_info completion callback can be called with a NULL skb, but
    only from net_zcopy_put and related API, used by zerocopy implementations
    that hold their own reference on the uarg, such as MSG_ZEROCOPY. This
    uarg is only ever completed from skb_zcopy_clear, so skb is always set.
    
    To prevent userspace from aliasing in-flight state on shared ring
    slots, allocate tpacket_uarg per packet, rather than per slot. This
    adds a small allocation to the transmit path. Use standard kmalloc to
    allow backporting to stable kernels.
    
    The uarg holds an sk_wmem_alloc reference, rather than an sk_refcnt
    reference. packet_free_tx_ring waits on sk_wmem_alloc before freeing
    the ring pages. Always allocate vec->deferred for tx_ring so page-backed
    rings also wait on sk_wmem_alloc when skb_copy_ubufs drops page refs
    before calling tpacket_ubuf_complete.
    
    Drop the tx_ring.pg_vec test that tpacket_destruct_skb performed before
    accessing the slot. The sk_wmem_alloc reference now guarantees that the
    slot is valid. The test is also not sufficient by itself, as it reads
    pg_vec without pg_vec_lock, so it can race with packet_set_ring.
    
    As a result a slot is released when its payload is copied, which can
    be before transmission (e.g., in skb_orphan_frags_rx). If copied
    before skb_tx_timestamp() is called, no slot timestamp is recorded,
    similar to when skb_orphan() was called early in the datapath before
    this patch.
    
    Revert the now unused previous skb_zcopy_.._nouarg infra.
    
    Depends on commit 992cc9f94ca9 ("net/packet: defer vmalloc TX_RING
    free until skbs finish").
    
    Reported-by: Katherine Leaver <[email protected]>
    Reported-by: Bjoern Doebel <[email protected]>
    Closes: https://lore.kernel.org/netdev/[email protected]/
    Fixes: 5cd8d46ea156 ("packet: copy user buffers before orphan or clone")
    Cc: [email protected]
    Signed-off-by: Willem de Bruijn <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
parisc: Increase kernel stack size to 32kb [+ + +]
Author: Helge Deller <[email protected]>
Date:   Sat Sep 19 22:09:20 2026 +0200

    parisc: Increase kernel stack size to 32kb
    
    commit 94b7e3a7e871ae27d4935c76959dfc61829f27f9 upstream.
    
    For 64-bit Linux kernels, increase the default kernel stack size
    (THREAD_SIZE_ORDER) to 32 kB, in order to avoid kernel crashes which have been
    triggered recently when building the debian vtk9 package with gcc 17:
    
    stackcheck: kworker/u128:0 will most likely overflow kernel stack (sp:179a83af0, stk bottom-top:179a80000-179a84000)
    Kernel panic - not syncing: low stack detected by irq handler - check messages
    CPU: 2 UID: 0 PID: 30760 Comm: kworker/u128:0 Tainted: G W 6.18.46-dirty #1 NONE
    Tainted: [W]=WARN
    Hardware name: 9000/800/rp3440
    Workqueue: writeback wb_workfn (flush-259:0)
    Backtrace:
     [<000000004022f050>] show_stack+0x70/0x90
     [<000000004022378c>] dump_stack_lvl+0x124/0x190
     [<000000004022382c>] dump_stack+0x34/0x48
     [<000000004020212c>] vpanic+0x204/0x648
     [<00000000402025c4>] panic+0x54/0x58
     [<0000000040232230>] do_cpu_irq_mask+0x3f8/0x440
     [<0000000040227070>] intr_return+0x0/0xc
    
    Signed-off-by: Helge Deller <[email protected]>
    Reported-by: John David Anglin <[email protected]>
    Cc: [email protected] # v6.18+
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

parisc: parse early parameters in setup_arch() [+ + +]
Author: Zhenghui Hao <[email protected]>
Date:   Fri Sep 18 13:46:49 2026 +0800

    parisc: parse early parameters in setup_arch()
    
    commit 289e99e7a263c6fd6a5d07d6d8f2156b70f3a9d5 upstream.
    
    parisc is one of the few architectures that does not call
    parse_early_param() from setup_arch().  That was mostly harmless until
    commit d49004c5f0c1 ("arch, mm: consolidate initialization of nodes,
    zones and memory map") moved the consumer of several hugetlb command
    line parameters into mm_core_init_early(), which runs before the
    generic parse_early_param() call in start_kernel().
    
    As a result hugepages=, hugepagesz=, default_hugepagesz=, hugetlb_cma=
    and hugetlb_free_vmemmap= are recorded after they have already been
    consumed and are silently dropped on parisc.
    
    Call parse_early_param() from setup_arch(), after the command line has
    been set up and the memory inventory has been taken.  jump_label_init()
    must be called first because early parameter handlers may enable or
    disable static keys.  Both functions are safe to call more than once:
    the generic calls in start_kernel() remain in place and turn into no-ops.
    
    Suggested-by: Mike Rapoport (Microsoft) <[email protected]>
    Fixes: d49004c5f0c1 ("arch, mm: consolidate initialization of nodes, zones and memory map")
    Cc: <[email protected]>
    Signed-off-by: Zhenghui Hao <[email protected]>
    Tested-by: Helge Deller <[email protected]>
    Signed-off-by: Helge Deller <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
PCI: Fix BAR resize for devices on a root bus [+ + +]
Author: Liz Fong-Jones <[email protected]>
Date:   Fri Sep 18 03:56:33 2026 +0000

    PCI: Fix BAR resize for devices on a root bus
    
    commit d58384c22739848efe14b34e9586e4f1242f33c0 upstream.
    
    pci_do_resource_release_and_resize() releases device BARs that share a
    bridge window with the BAR being resized, but when the device sits directly
    on a root bus (pdev->bus->self == NULL) it then skips resource assignment
    entirely and returns success, leaving the BARs it just released unassigned
    (IORESOURCE_UNSET).
    
    Skipping pbus_reassign_bridge_resources() is correct in that case -- there
    is no bridge window to adjust -- but the device BARs still have to be
    reassigned. Before the BAR release was consolidated into the PCI core, this
    case worked for amdgpu because the driver released the BARs itself and then
    called pci_assign_unassigned_bus_resources() unconditionally after the
    resize, which assigns unassigned device BARs also on a root bus. Commit
    db92e3fef53e ("drm/amdgpu: Remove driver side BAR release before resize")
    removed that call, so nothing assigns the released BARs anymore.
    
    This breaks amdgpu completely on the SolidRun HoneyComb LX2K (NXP LX2160A,
    arm64, ACPI), where ACPI doesn't expose the Root Port so the GPU endpoint
    appears directly on a "root bus" of its segment:
    
      amdgpu 0004:01:00.0: BAR 0 [mem 0xa400000000-0xa40fffffff 64bit pref]: releasing
      amdgpu 0004:01:00.0: BAR 2 [mem 0xa410000000-0xa4101fffff 64bit pref]: releasing
      amdgpu 0004:01:00.0: sw_init of IP block <gmc_v8_0> failed -19
      amdgpu 0004:01:00.0: amdgpu_device_ip_init failed
      amdgpu 0004:01:00.0: Fatal error during GPU init
    
    No error is logged because the resize path reports success; amdgpu then
    finds BAR 0 IORESOURCE_UNSET and bails out with -ENODEV.
    
    When there is no upstream bridge, call pci_bus_assign_resources() on the
    root bus to place the BARs released above, using the same alignment-sorted
    algorithm as normal enumeration instead of a manual per-BAR loop. This also
    walks the rest of the hierarchy under the root bus, as
    pci_assign_unassigned_bus_resources() used to for amdgpu before commit
    db92e3fef53e ("drm/amdgpu: Remove driver side BAR release before resize")
    removed that call -- the core-side fix that commit asked for ("such a
    problem should be fixed inside pci_resize_resource() instead").
    
    pci_bus_assign_resources() returns void, so failure is detected by checking
    whether the released BARs are still assigned afterward; if not, roll back
    as in the bridged case. This is stricter than the bridged path -- it fails
    on any unplaced resource, not just required ones -- since a root bus
    typically has one shared window, and failing loudly seemed better than
    leaving something silently unassigned.
    
    The root bus path also had a locking bug that any fix here necessarily
    touches: the old "goto out" jumped to up_read(&pci_bus_sem) without a
    matching down_read() (as does the "goto restore" taken when
    pci_dev_res_add_to_list() fails in the release loop). Take pci_bus_sem
    before the BAR release loop so every path through the function holds it
    exactly once.
    
    Fixes: 337b1b566db0 ("PCI: Fix restoring BARs on BAR resize rollback path")
    Link: https://bugs.launchpad.net/ubuntu/+source/linux-hwe-7.0/+bug/2159596
    Suggested-by: Ilpo Järvinen <[email protected]>
    Assisted-by: Claude:claude-fable-5 checkpatch
    Assisted-by: Claude:claude-sonnet-5
    Signed-off-by: Liz Fong-Jones <[email protected]>
    [bhelgaas: commit log]
    Signed-off-by: Bjorn Helgaas <[email protected]>
    Reviewed-by: Ilpo Järvinen <[email protected]>
    Cc: [email protected]
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

PCI: of_property: Omit bus properties without a subordinate bus [+ + +]
Author: Angel J <[email protected]>
Date:   Fri Sep 18 14:55:40 2026 -0500

    PCI: of_property: Omit bus properties without a subordinate bus
    
    commit 8805840aad73df7146778be243a196d48b4f6430 upstream.
    
    A bridge (a device with a Type 1 header) may not have a secondary bus
    allocated (pdev->subordinate), e.g., if there are no available bus numbers
    or the bridge secondary/subordinate bus numbers are not writable.
    
    The dynamic OF helpers of_pci_prop_bus_range() and of_pci_prop_intr_map()
    dereference pdev->subordinate without checking it.  When
    CONFIG_PCI_DYNAMIC_OF_NODES is enabled, this can cause a NULL pointer
    dereference and early boot hang.
    
    Generate 'bus-range' and 'interrupt-map' properties only when a subordinate
    bus exists.  Keep the node and its remaining properties for bridges without
    one.
    
    The problem was latent since 407d1a51921e ("PCI: Create device tree node
    for bridge"), but wasn't reachable until 1f340724419e ("PCI: of: Create
    device tree PCI host bridge node"), which appeared in v6.15.  Before
    1f340724419e, of_pci_make_dev_node() returned early because the parent OF
    node was missing.
    
    Fixes: 407d1a51921e ("PCI: Create device tree node for bridge")
    Signed-off-by: Angel J <[email protected]>
    [bhelgaas: move pdev->subordinate test to callees, commit log]
    Signed-off-by: Bjorn Helgaas <[email protected]>
    Cc: [email protected]      # v6.6+
    Link: https://patch.msgid.link/20260918195540.GA1187209@bhelgaas
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
perf/core: Fill branch entries with a single assignment [+ + +]
Author: Puranjay Mohan <[email protected]>
Date:   Mon Aug 10 06:35:36 2026 -0700

    perf/core: Fill branch entries with a single assignment
    
    [ Upstream commit 24b620729e53d978b3e425f55bc66efd3bab1f59 ]
    
    perf_clear_branch_entry_bitfields() clears the bitfields of struct
    perf_branch_entry one by one and leaves from/to alone, since callers
    overwrite those straight away. The list has to be kept in sync with the
    struct by hand and has already fallen behind: new_type and priv were
    added to perf_branch_entry and never added here.
    
    Only BRBE writes those two, and neither for every record.
    brbe_set_perf_entry_type() leaves new_type alone for a branch type it
    does not recognise, and priv is not set for source-only records.
    arm_pmuv3.c allocates the per-CPU branch stack with kmalloc(), so such a
    record reaches userspace with whatever the slot held: uninitialised
    kmalloc() data on the first pass over the buffer, the previous record's
    values after that. Nothing under arch/x86/events/ writes either field,
    so only arm64 is affected.
    
    Assign the whole entry at each site instead. Everything not named is
    then zero, and there is no list to keep in sync. The bitfields add up to
    exactly 64 bits, so the struct has no padding to leave undefined.
    
    perf_clear_branch_entry_bitfields() has no callers left, so remove it.
    perf_entry_from_brbe_regset() assigns an empty literal instead, since it
    fills from/to conditionally. PERF_BR_SPEC_NA is 0, so dropping the
    explicit spec assignment changes nothing.
    
    Fixes: b190bc4ac9e6 ("perf: Extend branch type classification")
    Fixes: 5402d25aa571 ("perf: Capture branch privilege information")
    Suggested-by: Peter Zijlstra <[email protected]>
    Signed-off-by: Puranjay Mohan <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Tested-by: Yifan Wu <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

perf/core: Fix a refcount leak in attach_perf_ctx_data() [+ + +]
Author: Namhyung Kim <[email protected]>
Date:   Sun Sep 20 16:16:39 2026 -0700

    perf/core: Fix a refcount leak in attach_perf_ctx_data()
    
    [ Upstream commit cca4980630b3c7a85f53cb43c6018184ce5d4e37 ]
    
    The attach_perf_ctx_data() can race on global and !global cases.  The
    global case is protected by global_ctx_data_rwsem and shares a single
    reference count using perf_ctx_data.global field.
    
    But when it races with !global case, it may miss to set the global field
    and result in a reference count leak.
    
     CPU1                                  CPU2
     ----------------------------------------------------------------
     attach_task_ctx_data(.global=1)       attach_task_ctx_data(.global=0)
       cd1 = alloc_perf_ctx_data();          cd2 = alloc_perf_ctx_data();
                                             //    { .global = 0, .refcount = 1 };
    
                                             try_cmpxchg(); // success,
                                             // task->perf_ctx_data = cd2
       try_cmpxhg(); // fail; old = cd2
       refcount_inc_not_zero(&old->refcount); // success
         // old.refcount = 2
       free_perf_ctx_data(cd1);
    
    Then later detach_global_ctx_data() will see the data but it's not
    marked as global, so it won't call detach_task_ctx_data().
    
    Fixes: 506e64e710ff ("perf: attach/detach PMU specific data")
    Assisted-by: Sashiko.dev:Gemini-3.1-pro
    Signed-off-by: Namhyung Kim <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

perf/core: Fix NULL pmu_ctx passed to pmu->sched_task() [+ + +]
Author: Puranjay Mohan <[email protected]>
Date:   Mon Aug 10 06:35:34 2026 -0700

    perf/core: Fix NULL pmu_ctx passed to pmu->sched_task()
    
    commit 36bb85cf36cab15fb611cb44b78a5df06e4e69a2 upstream.
    
    perf_pmu_sched_task() returns early when cpuctx->task_ctx is set, and
    cpc->task_epc is only non-NULL while a task context is scheduled in on
    this CPU. __perf_pmu_sched_task() therefore always passes NULL:
    
      Unable to handle kernel NULL pointer dereference at virtual address 00
      pc : armv8pmu_sched_task+0x14/0x50
      Call trace:
       armv8pmu_sched_task+0x14/0x50 (P)
       perf_pmu_sched_task+0xac/0x108
       __perf_event_task_sched_out+0x6c/0xe0
    
    Pass &cpc->epc instead, the CPU-wide context for this PMU, which the
    function already dereferences a few lines up to find pmu.
    
    armv8pmu_sched_task() is the only in-tree implementation that
    dereferences the argument, and it only reads ->pmu, so the oops needs
    BRBE, added in v6.17.
    
    Fixes: bd2756811766 ("perf: Rewrite core context handling")
    Signed-off-by: Puranjay Mohan <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Tested-by: Yifan Wu <[email protected]>
    Cc: [email protected]
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

perf/core: Run sched_task() for PMUs with only CPU-wide events [+ + +]
Author: Puranjay Mohan <[email protected]>
Date:   Mon Aug 10 06:35:35 2026 -0700

    perf/core: Run sched_task() for PMUs with only CPU-wide events
    
    commit 3d8d74100954a3b17e5c5e37adfe14e16b1db103 upstream.
    
    perf_pmu_sched_task() returns early when cpuctx->task_ctx is set and
    leaves the work to perf_ctx_sched_task_cb(), which only walks
    ctx->pmu_ctx_list. A PMU whose events are all CPU-wide is not on that
    list, so nothing calls its sched_task(). With
    
      perf record -b -e cycles -a -- ls
    
    armv8pmu_sched_task() is skipped on every switch to a task that has a
    perf context but no event on that PMU, and BRBE records leak across the
    task boundary. intel_pmu_lbr_add() calls perf_sched_cb_inc()
    unconditionally too, so LBR records leak the same way on x86.
    
    Drop the early return and skip only the CPCs that
    perf_ctx_sched_task_cb() handles. That one needs a gate of its own to
    make the split exact: it tests cpc->sched_cb_usage, which
    perf_sched_cb_inc() sets per CPU for every branch stack user, so a task
    with an event for that PMU pinned to another CPU would be handled twice.
    On x86 the second __intel_pmu_lbr_restore() finds lbr_stack_state ==
    LBR_NONE and calls intel_pmu_lbr_reset(), throwing away the callstack
    the first one restored.
    
    cpc->task_epc is set only while a task context is scheduled in, and
    there is one epc per PMU on ctx->pmu_ctx_list, so the two gates are
    inverses.
    
    For the CPCs perf_pmu_sched_task() picks up, the callback now runs
    outside the perf_ctx_disable() and perf_ctx_enable() pair in
    perf_event_context_sched_in(). __perf_pmu_sched_task() disables the PMU
    around the call itself.
    
    Fixes: bd2756811766 ("perf: Rewrite core context handling")
    Signed-off-by: Puranjay Mohan <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Tested-by: Yifan Wu <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Cc: [email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3 [+ + +]
Author: Dapeng Mi <[email protected]>
Date:   Thu Sep 17 09:52:31 2026 +0800

    perf/x86/intel: Constrain Panther Cove UOPS_DISPATCHED events to PMCs 0-3
    
    commit 04a7ef3b7aa3af34202285043e3423a2e3983595 upstream.
    
    Per the latest Panther Cove event definitions, the following events are
    only supported on PMCs 0-3:
    
    - UOPS_DISPATCHED.INT_EU_ALL (0x1b2)
    - UOPS_DISPATCHED.ALU (0x2b2)
    
    Add explicit event constraints for these two events so scheduling does
    not place them on unsupported counters.
    
    Fixes: d345b6bb8860 ("perf/x86/intel: Add core PMU support for DMR")
    Signed-off-by: Dapeng Mi <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Cc: <[email protected]> # v7.0+
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

perf/x86/intel: Delete dead NVL PEBS data-source initcall [+ + +]
Author: Dapeng Mi <[email protected]>
Date:   Thu Sep 17 09:52:30 2026 +0800

    perf/x86/intel: Delete dead NVL PEBS data-source initcall
    
    [ Upstream commit 858b37ca19d3f695f7fa94cd14be5a856fd8e7d7 ]
    
    Nova Lake now uses the OMR data-source table for PEBS data-source
    decoding and no longer depends on the legacy static pebs_data_source[]
    mapping.
    
    Remove the dead intel_pmu_pebs_data_source_lnl() initialization call
    for NVL.
    
    Fixes: c847a208f43b ("perf/x86/intel: Add core PMU support for Novalake")
    Signed-off-by: Dapeng Mi <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused [+ + +]
Author: Sean Christopherson <[email protected]>
Date:   Mon Sep 21 12:14:11 2026 -0700

    perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused
    
    [ Upstream commit d06260e99eb93d2942b7af4ccd789eb8a6c829d3 ]
    
    When filling the list of MSRs to be loaded by KVM on VM-Enter and VM-Exit,
    load the guest values for DS_AREA and (conditionally) MSR_PEBS_DATA_CFG if
    and only if PEBS will be active in the guest, i.e. only if a PEBS record
    may be generated while running the guest.  As shown by the !pebs_ept path,
    it's perfectly safe to run with the host's DS_AREA, so long as PEBS-enabled
    counters are disabled via PERF_GLOBAL_CTRL.
    
    Omitting DS_AREA and MSR_PEBS_DATA_CFG when PEBS is unused saves two MSR
    writes per MSR on each VMX transition, i.e. eliminates two/four pointless
    MSR writes on each VMX roundtrip when PEBS isn't being used by the guest.
    
    Fixes: c59a1f106f5c ("KVM: x86/pmu: Add IA32_PEBS_ENABLE MSR emulation for extended PEBS")
    Signed-off-by: Sean Christopherson <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Reviewed-by: Jim Mattson <[email protected]>
    Reviewed-by: Dapeng Mi <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED [+ + +]
Author: Sean Christopherson <[email protected]>
Date:   Mon Sep 21 12:14:10 2026 -0700

    perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED
    
    [ Upstream commit 4b64dbdc5861477f148e13d1ed127e7fe7182e4f ]
    
    When filling the list of MSRs to be loaded by KVM on VM-Enter and VM-Exit,
    *never* insert an entry for PEBS_ENABLED if the CPU properly isolates PEBS
    events, in which case disabling counters via PERF_GLOBAL_CTRL is sufficient
    to prevent unwanted PEBS events in the guest (or host).  Because perf loads
    PEBS_ENABLE with the unfiltered cpu_hw_events.pebs_enabled, i.e. with both
    host and guest masks, there is no need to load different values for the
    guest versus host, perf+KVM can and should simply control which counters
    are enabled/disabled via PERF_GLOBAL_CTRL.
    
    Avoiding touching PEBS_ENABLED "fixes" a bug where PEBS_ENABLED can end up
    with "stuck" bits if a PEBS event is throttled between generating the list
    and actually entering the guest (Intel CPUs can't arbtitrarily block NMIs).
    Fixes in quotes because leaving PEBS_ENABLED as-is doesn't fix the
    underlying problem of perf (via PMIs) being able to modify state after the
    perf<=>KVM handoff.
    
    But not writing PEBS_ENABLED is desirable no matter what, as stating the
    obvious, leaving PEBS_ENABLED as-is avoids three MSR writes on every VMX
    transition: one each on entry/exit, and one more explicit WRMSR to zero
    PEBS_ENABLED before VM-Entry (KVM assumes the only reason PEBS_ENABLED is
    in the load list is if the CPU lacks PEBS isolation and thus needs a
    quiescent period).
    
    Opportunistically add comments to (better) explain the rules for generating
    the set of PEBS counters that will be active while the guest is running,
    along with a FIXME for the suspected hack-a-fix where perf disables guest
    PEBS if _any_ PEBS event is configured to count in the host (commit
    854250329c02 ("KVM: x86/pmu: Disable guest PEBS temporarily in two rare
    situations") doesn't explain the motivation, at all).
    
    Fixes: c59a1f106f5c ("KVM: x86/pmu: Add IA32_PEBS_ENABLE MSR emulation for extended PEBS")
    Signed-off-by: Sean Christopherson <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Reviewed-by: Dapeng Mi <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits [+ + +]
Author: Sean Christopherson <[email protected]>
Date:   Mon Sep 21 12:14:09 2026 -0700

    perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits
    
    [ Upstream commit cec38d5c098a350dcf084d345025136ade7e6d1e ]
    
    When reinstating PEBS counters into PERF_GLOBAL_CTRL for a KVM guest, mask
    the value with perf's desired/original PERF_GLOBAL_CTRL value to ensure
    KVM doesn't unintentionally set reserved bits in PERF_GLOBAL_CTRL.  E.g.
    if the guest's PEBS_ENABLE value had bit 63, "Enable Precise Store", set,
    then using the raw guest PEBS value would propagate bit 63 to the guest's
    PERF_GLOBAL_CTRL value (which thankfully would be a failed VM-Entry, not
    a VMX Abort).
    
    The only reason this bug isn't reachable is because KVM doesn't support
    "Enable Precise Store" (which is probably a KVM bug?), i.e. bit 63 can't
    be set in kvm_pmu->pebs_enable and thus not in arr[pebs_enable].guest.  In
    other words, this _should_ be a glorified NOP in the current code base.
    
    Fixes: c59a1f106f5c ("KVM: x86/pmu: Add IA32_PEBS_ENABLE MSR emulation for extended PEBS")
    Signed-off-by: Sean Christopherson <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Reviewed-by: Dapeng Mi <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification [+ + +]
Author: Dapeng Mi <[email protected]>
Date:   Thu Sep 17 09:52:24 2026 +0800

    perf/x86/intel: Fix CMT PEBS load/store direction for latency events, to fix sample classification
    
    commit e961d6db42d1f1b66c81438969a648f08573bd9c upstream.
    
    The same bug exists on Crestmont as on Gracemont:
    intel_cmt_pebs_event_constraints[] applies LAT_CONSTRAINT constraints to
    MEM_UOPS_RETIRED.{LOAD,STORE}_LATENCY, but does not set explicit
    LOAD/STORE flags for those events.
    
    The PEBS latency path (pebs_latency_data(), via cmt_latency_data) uses
    the event flags to determine memory operation direction. Without an
    explicit STORE flag, samples from MEM_UOPS_RETIRED.STORE_LATENCY can be
    misclassified as LOADs.
    
    Set explicit LOAD/STORE flags in intel_cmt_pebs_event_constraints[] for:
    
    - MEM_UOPS_RETIRED.LOAD_LATENCY
    - MEM_UOPS_RETIRED.STORE_LATENCY
    
    This fixes incorrect STORE sample classification.
    
    Fixes: e99fb45436ea ("perf/x86/intel: Update event constraints and cache_extra_regsfor MTL")
    Signed-off-by: Dapeng Mi <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Cc: <[email protected]> # v7.2+
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification [+ + +]
Author: Dapeng Mi <[email protected]>
Date:   Thu Sep 17 09:52:25 2026 +0800

    perf/x86/intel: Fix DKT PEBS load/store direction for latency events, to fix sample classification
    
    commit 8302c5f475fa5a4ed36a7d32c7b965023d5d7066 upstream.
    
    Same bug exists on Darkmont as on Gracemont:
    intel_dkt_pebs_event_constraints[] applies LAT_CONSTRAINT constraints to
    MEM_UOPS_RETIRED.{LOAD,STORE}_LATENCY, but does not set explicit
    LOAD/STORE flags for those events.
    
    The PEBS latency path (pebs_latency_data(), via cmt_latency_data) uses
    the event flags to determine memory operation direction. Without an
    explicit STORE flag, samples from MEM_UOPS_RETIRED.STORE_LATENCY can be
    misclassified as LOADs.
    
    Set explicit LOAD/STORE flags in intel_dkt_pebs_event_constraints[] for:
    
    - MEM_UOPS_RETIRED.LOAD_LATENCY
    - MEM_UOPS_RETIRED.STORE_LATENCY
    
    This fixes incorrect STORE sample classification. Additionally remove
    INTEL_HYBRID_LAT_CONSTRAINT() since no one uses it anymore.
    
    Fixes: 65fd435095bb ("perf/x86/intel: Update event constraints for PTL")
    Signed-off-by: Dapeng Mi <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Cc: <[email protected]> # v7.2+
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification [+ + +]
Author: Dapeng Mi <[email protected]>
Date:   Thu Sep 17 09:52:23 2026 +0800

    perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification
    
    commit 89dc568e8c0be60e05e5fcd0b528c79077d7b84b upstream.
    
    On Gracemont, intel_grt_pebs_event_constraints[] applies LAT_CONSTRAINT
    constraints to MEM_UOPS_RETIRED.{LOAD,STORE}_LATENCY, but does not set
    explicit LOAD/STORE flags for those events.
    
    The PEBS latency path (pebs_latency_data(), via __grt_latency_data())
    uses the event flags to determine memory operation direction. Without an
    explicit STORE flag, samples from MEM_UOPS_RETIRED.STORE_LATENCY can be
    misclassified as LOADs.
    
    Set explicit LOAD/STORE flags in intel_grt_pebs_event_constraints[] for:
    
    - MEM_UOPS_RETIRED.LOAD_LATENCY
    - MEM_UOPS_RETIRED.STORE_LATENCY
    
    Also update __grt_latency_data() to explicitly interpret these flags when
    assigning the sampled memory operation direction.
    
    This fixes incorrect STORE sample classification.
    
    Fixes: 39a41278f041 ("perf/x86/intel: Fix PEBS memory access info encoding for ADL")
    Signed-off-by: Dapeng Mi <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Cc: <[email protected]> # v7.2+
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

perf/x86/intel: Fix Panther Cove PEBS data-source snoop states [+ + +]
Author: Dapeng Mi <[email protected]>
Date:   Thu Sep 17 09:52:29 2026 +0800

    perf/x86/intel: Fix Panther Cove PEBS data-source snoop states
    
    commit 0ac5d6ca2c3b7d047963b67dccef17902dc6c017 upstream.
    
    For Panther Cove, the snoop states for the data source encodings
    "Prefetch Promotion" and "Cross Core Prefetch Promotion" should be
    SNOOP_NONE instead of SNOOP_MISS.
    
    Fix the incorrect snooping states for Panther Cove.
    
    Fixes: d2bdcde9626c ("perf/x86/intel: Add support for PEBS memory auxiliary info field in DMR")
    Signed-off-by: Dapeng Mi <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Cc: <[email protected]> # v7.0+
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs() [+ + +]
Author: Sean Christopherson <[email protected]>
Date:   Mon Sep 21 12:14:12 2026 -0700

    perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs()
    
    [ Upstream commit a391618e1d563f099e4c2a704f45d08329ccdf7c ]
    
    Drop "support" for passing a NULL @data/@kvm_pmu param when getting guest
    MSRs.  KVM, the only in-tree user, unconditionally passes a non-NULL
    pointer, and carrying code that suggests @data may be NULL is confusing,
    e.g. incorrectly implies that there are scenarios where KVM doesn't pass
    a PMU context.
    
    Fixes: 8183a538cd95 ("KVM: x86/pmu: Add IA32_DS_AREA MSR emulation to support guest DS")
    Signed-off-by: Sean Christopherson <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Reviewed-by: Jim Mattson <[email protected]>
    Reviewed-by: Dapeng Mi <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints [+ + +]
Author: Dapeng Mi <[email protected]>
Date:   Thu Sep 17 09:52:27 2026 +0800

    perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints
    
    commit 7c944595cc43a664019edc22505b6d6f6039be5e upstream.
    
    On Lion Cove, PEBS data source is valid only for these events:
    
    - MEM_TRANS_RETIRED.LOAD_LATENCY (0x1cd)
    - MEM_TRANS_RETIRED.STORE_SAMPLE (0x2cd)
    
    The perfmon database (https://github.com/intel/perfmon) previously
    tagged additional memory events such as MEM_INST_RETIRED.STLB_MISS_LOADS
    with L1_Hit_Indication, implying PEBS data-source support, which is
    incorrect. The database has since been fixed, but
    intel_lnc_pebs_event_constraints[] still follows the old definition and
    marks those events as data-source capable.
    
    As a result, get_data_src() may decode data-source information for
    events that do not provide valid PEBS data-source data and mislead
    users.
    
    Remove those non-data-source memory events from the Lion Cove PEBS
    constraint table so matching falls back to the regular non-PEBS
    constraints, which already provide the same counter constraints.
    
    Also update lnc_latency_data() to decode LOAD/STORE flags explicitly
    when setting memory operation direction, for consistency with other
    *_latency_data() helpers.
    
    Fixes: a932aa0e868f ("perf/x86: Add Lunar Lake and Arrow Lake support")
    Signed-off-by: Dapeng Mi <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Cc: <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints [+ + +]
Author: Dapeng Mi <[email protected]>
Date:   Thu Sep 17 09:52:28 2026 +0800

    perf/x86/intel: Remove incorrect Panther Cove PEBS data-source constraints
    
    commit 9f93d33ad65af9853974c1ec4ba1f937e8ba1376 upstream.
    
    Same issue exists on Panther Cove, PEBS data source is valid only for
    these events:
    
    - MEM_TRANS_RETIRED.LOAD_LATENCY (0x1cd)
    - MEM_TRANS_RETIRED.STORE_SAMPLE (0x2cd)
    
    The perfmon database (https://github.com/intel/perfmon) previously
    tagged additional memory events such as MEM_INST_RETIRED.STLB_MISS_LOADS
    with L1_Hit_Indication, implying PEBS data-source support, which is
    incorrect. The database has since been fixed, but
    intel_pnc_pebs_event_constraints[] still follows the old definition and
    marks those events as data-source capable.
    
    As a result, get_data_src() may decode data-source information for
    events that do not provide valid PEBS data-source data and mislead
    users.
    
    Remove those non-data-source memory events from the Pather Cove PEBS
    constraint table so matching falls back to the regular non-PEBS
    constraints, which already provide the same counter constraints.
    
    Also update pnc_latency_data() to decode LOAD/STORE flags explicitly
    when setting memory operation direction, for consistency with other
    *_latency_data() helpers.
    
    Fixes: d345b6bb8860 ("perf/x86/intel: Add core PMU support for DMR")
    Signed-off-by: Dapeng Mi <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Cc: <[email protected]> # v7.0+
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
pinctrl: generic: serialise pinctrl_generic_dt_node_to_map() [+ + +]
Author: Sarah Emery <[email protected]>
Date:   Fri Aug 28 17:58:17 2026 +0200

    pinctrl: generic: serialise pinctrl_generic_dt_node_to_map()
    
    [ Upstream commit 51ae99659469edba2e83931efa26b77d36ca02c2 ]
    
    pinctrl_generic_add_group() documents that the caller must take care of
    locking, and pinmux_generic_add_function() needs it too, but
    pinctrl_generic_dt_node_to_map() calls them without holding
    pctldev->mutex, and the core caller in create_pinctrl() does not take it
    either.
    
    The driver core calls pinctrl_bind_pins() before probing a device, so
    two devices that reference the same pin controller can run
    pinctrl_generic_dt_node_to_map() on one pctldev at the same time.
    
    Both `add` functions take the new selector from pctldev->num_groups or
    pctldev->num_functions, and radix_tree_insert() at that index.
    Two racing callers can read the same selector before either
    has inserted, so the second insert collides and fails:
    
      k1-pinctrl d401e000.pinctrl:
      error -EEXIST: error adding function pcie2-0-cfg
      k1-pinctrl d401e000.pinctrl:
      does not have pin group pcie0-0-cfg.pcie0-0-pins
    
    leaving one consumer without its pin configuration.
    
    This was hit on a SpacemiT K3 board, where PCIe devices probe in parallel
    against the single shared pin controller.
    
    Take pctldev->mutex across the whole function, so that the groups and the
    function referring are in a single critical section.
    
    Fixes: 43722575e5cd ("pinctrl: add generic functions + pins mapper")
    Signed-off-by: Sarah Emery <[email protected]>
    Signed-off-by: Linus Walleij <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

pinctrl: meson: Fix typo in s4 group name [+ + +]
Author: Sean Anderson <[email protected]>
Date:   Mon Aug 17 12:22:11 2026 -0400

    pinctrl: meson: Fix typo in s4 group name
    
    [ Upstream commit 692f32609a30f75ca3401e25b504bfd06bd5662a ]
    
    One of the i2c pin groups has some junk at the end. The name should be
    i2c2_scl_h1, and indeed that's the name used by i2c2_pins3 in
    meson-s4.dtsi.
    
    Fixes: 775214d389c25 ("pinctrl: meson: add pinctrl driver support for Meson-S4 Soc")
    Signed-off-by: Sean Anderson <[email protected]>
    Reviewed-by: Neil Armstrong <[email protected]>
    Signed-off-by: Linus Walleij <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

pinctrl: mpfs-mssio: fix width of unused bank voltage setting [+ + +]
Author: Conor Dooley <[email protected]>
Date:   Mon Aug 31 11:50:58 2026 +0100

    pinctrl: mpfs-mssio: fix width of unused bank voltage setting
    
    commit e6c5d6c589764a41218cbd81bd7d6c9b906e2cf0 upstream.
    
    The bank voltages are only 4 bits wide, so when a pin was unused the
    driver was not correctly interpreting it as being at zero volts, because
    the driver's value for unused had two extra set bits.
    
    CC: [email protected]
    Fixes: 488d704ed7b7 ("pinctrl: add polarfire soc mssio pinctrl driver")
    Signed-off-by: Conor Dooley <[email protected]>
    Signed-off-by: Linus Walleij <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

pinctrl: mpfs-mssio: use correct regmap function to set bank voltage [+ + +]
Author: Conor Dooley <[email protected]>
Date:   Mon Aug 31 11:50:59 2026 +0100

    pinctrl: mpfs-mssio: use correct regmap function to set bank voltage
    
    commit 8039b75af5808572ff3656eaed07d348aed3989a upstream.
    
    regmap_assign_bits() is not the correct function to use for an RMW
    operation, as it maps to regmap_set_bits() or regmap_clear_bits() and
    the former will never zero a bit. Use regmap_update_bits() instead,
    which will actually set the bank voltages to what have been requested.
    
    CC: [email protected]
    Fixes: 488d704ed7b7 ("pinctrl: add polarfire soc mssio pinctrl driver")
    Signed-off-by: Conor Dooley <[email protected]>
    Signed-off-by: Linus Walleij <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

pinctrl: qcom: ipq5210: Publish the OF module alias [+ + +]
Author: hpp.iscas <[email protected]>
Date:   Sat Sep 5 21:43:40 2026 +0800

    pinctrl: qcom: ipq5210: Publish the OF module alias
    
    [ Upstream commit 447dcff90557748a5ac384aaf5d70b2c7e3b961d ]
    
    The IPQ5210 TLMM platform driver can be a module and matches through
    ipq5210_tlmm_of_match. This table is not published for OF modalias
    matching.
    
    Publish the existing table, preserving arch_initcall ordering and
    the shared MSM pinctrl probe.
    
    Fixes: a549fe22376f ("pinctrl: qcom: Introduce IPQ5210 TLMM driver")
    Signed-off-by: hpp.iscas <[email protected]>
    Reviewed-by: Konrad Dybcio <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Bartosz Golaszewski <[email protected]>
    Signed-off-by: Linus Walleij <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

pinctrl: qcom: nord: Split QUP1 SE2/SE3 into lane-pair functions [+ + +]
Author: Shawn Guo <[email protected]>
Date:   Sun Aug 30 11:02:01 2026 +0800

    pinctrl: qcom: nord: Split QUP1 SE2/SE3 into lane-pair functions
    
    commit b0be01f133aa36d2fd12d71c916cb8a47826b4ed upstream.
    
    QUP1 SE2 and SE3 pack all four of their lanes pair-wise onto only two
    pins each: lanes 0/1 (I2C SDA/SCL) at mux value 2 and lanes 2/3 (UART
    TX/RX) at mux value 1, on gpio127/gpio128 and gpio129/gpio130
    respectively.
    
    Both mux values were named "qup1_se2" (respectively "qup1_se3"), so the
    two distinct lane pairs became indistinguishable. msm_pinmux_set_mux()
    stops at the first entry matching the requested function, which means
    mux value 1 was always selected and the I2C lanes could never be muxed
    out. In practice i2c9 and i2c10 got the UART lanes and did not work,
    while uart9 and uart10 happened to be muxed correctly.
    
    Give each lane pair its own function, following the _01/_23 naming
    already used for the same hardware arrangement by the shikra, eliza,
    hawi and maili TLMM drivers. Both functions still cover the full pin
    pair, so a single pinctrl state per protocol remains sufficient.
    
    Drop gpio129/gpio130 from the SE2 group list, since those pins
    belong to SE3 and were never reachable through the SE2 function.
    
    Also rename QUP1 SE2/SE3 functions in the binding doc accordingly.
    
    While at it, add missing "gpio", "qup3_se0_mira" and "qup3_se0_mirb"
    to the binding function enum to get the list complete.
    
    Fixes: c24dd0826f06 ("pinctrl: qcom: add the TLMM driver for the Nord platforms")
    Cc: [email protected]
    Assisted-by: LLM
    Signed-off-by: Shawn Guo <[email protected]>
    Reviewed-by: Konrad Dybcio <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Bartosz Golaszewski <[email protected]>
    Signed-off-by: Linus Walleij <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

pinctrl: single: free the IRQ on domain creation failure [+ + +]
Author: Myeonghun Pak <[email protected]>
Date:   Sun Sep 13 00:03:12 2026 -0400

    pinctrl: single: free the IRQ on domain creation failure
    
    commit 1d9bb9c870632adea249cfdec49ac37e6f069164 upstream.
    
    pcs_irq_init_chained_handler() requests a shared IRQ on affected SoCs, but
    its domain creation failure path only removes a chained handler. That does
    not release the action installed by request_irq(). The probe can continue
    without interrupt support while leaving the shared IRQ action registered.
    
    Use pcs_irq_free() to undo the appropriate type of handler registration.
    At this point pcs->domain is NULL, so the helper only releases the parent
    IRQ handler. Then mark the IRQ invalid, as the other initialization error
    paths already do, to prevent another release from a later probe unwind or
    remove.
    
    This issue was identified during our ongoing static-analysis research while
    reviewing kernel code.
    
    Fixes: 3e6cee1786a1 ("pinctrl: single: Add support for wake-up interrupts")
    Cc: [email protected]
    Assisted-by: LLM
    Co-developed-by: Ijae Kim <[email protected]>
    Signed-off-by: Ijae Kim <[email protected]>
    Signed-off-by: Myeonghun Pak <[email protected]>
    Signed-off-by: Linus Walleij <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

pinctrl: sunxi: A523: fix voltage withstand encoding [+ + +]
Author: Andre Przywara <[email protected]>
Date:   Mon Sep 14 12:07:44 2026 +0200

    pinctrl: sunxi: A523: fix voltage withstand encoding
    
    [ Upstream commit ef56085dfd1df3a53ecce8a4ff440cf6d67f430d ]
    
    The Allwinner A523 uses the same GPIO voltage "withstand" programming
    (setting the input level voltage thresholds) as the previous SoCs, but
    for some odd reason inverts the encoding of 1.8V vs. 3.3V.
    
    Add a new bias voltage type to note this difference, and select it for
    the A523. At the same time also use the newer "CTL" version, which in
    addition allows to turn off the withstand programming for I/O voltages
    other than exact 1.8V or 3.3V (for instance for 2.5V sometimes used for
    Ethernet PHYs). The A523 has that enable register, but didn't use it
    so far.
    
    This fixes eMMC and reportedly Ethernet operation on some A523 boards.
    
    Fixes: 648be4cd9517 ("pinctrl: sunxi: Add support for the Allwinner A523")
    Signed-off-by: Andre Przywara <[email protected]>
    Tested-by: Per Larsson <[email protected]>
    Tested-by: Juan Manuel Lopez Carrillo <[email protected]>
    Reviewed-by: Chen-Yu Tsai <[email protected]>
    Tested-by: Chen-Yu Tsai <[email protected]> # Fixes eMMC on Orange Pi 4A
    Signed-off-by: Linus Walleij <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

pinctrl: sunxi: keep a shadow copy of the data register output latches [+ + +]
Author: Ilya Titov <[email protected]>
Date:   Thu Sep 3 12:17:18 2026 +0300

    pinctrl: sunxi: keep a shadow copy of the data register output latches
    
    commit a13f7f5d14af9baf34eb12c25962f8d3542b281d upstream.
    
    On Allwinner SoCs, reading a bank's data register returns the pin level,
    not the output latch, for pins that are muxed as inputs.  Writing a GPIO
    therefore corrupts the output latches of all input-muxed pins in the
    same bank: the read-modify-write in sunxi_pinctrl_gpio_set() reads back
    their pin levels and writes those into their latches.
    
    This breaks emulated open-drain lines (e.g. a bit-banged I2C bus from
    i2c-gpio).  Such a line is released high by muxing it as input and
    letting the pull-up raise it, so any concurrent GPIO write in the same
    bank stores 1 into its latch.  Driving the line low afterwards is a
    non-atomic data-then-mux sequence in sunxi_pinctrl_gpio_direction_output();
    if the poisoning write lands between the two steps, the pin actively
    drives high (push-pull) instead of low.
    
    Observed in practice as sporadic glitches on a T507 board bit-banging
    I2C on port E while other PE GPIOs are toggled.  On a scope the failure
    is unmistakable: on a clock pulse where SCL should fall to GND, the line
    instead steps *above* its idle high level for the whole low phase — the
    pad drives a strong push-pull 3.3 V high, higher than the level the
    pull-up sustains on the loaded bus — before the next transition recovers
    it.  The same can hit SDA, corrupting data instead of clocks.
    
    Steps to reproduce on any sunxi board with a bit-banged (i2c-gpio) bus:
    
      # background: toggle any other GPIO of the same bank, e.g. line 21
      gpioset -c <chip> --toggle 100us 21=0 &
    
      # foreground: keep the bit-banged bus busy
      while :; do i2cdetect -y <bus> 0x50 0x57; done
    
      # watch SCL/SDA with a scope or logic analyzer: sporadic clock-low
      # phases driven high (above the pull-up level) instead of low
    
    The bank spinlock cannot help: the racing write is a perfectly valid
    whole-register RMW that faithfully writes back what the hardware
    returned.  There are no set/clear registers on this IP to write a single
    bit atomically.
    
    Fix it the same way gpio-mmio handles hardware whose data register read
    does not return the output latch: keep a shadow copy of each bank's
    latches, base the read-modify-write on the shadow, and only write the
    register.  The shadow is seeded from the hardware at probe time so pins
    left in output mode by the bootloader keep their state.  Pins that reach
    output mode through the gpiolib paths write their value (and thereby
    their shadow bit) before the mux switch in
    sunxi_pinctrl_gpio_direction_output(); pins muxed to gpio_out directly
    through a pinmux node bypass that path, so sunxi_pmx_set() refreshes
    their shadow bit from the latch (readable once the pin is in output
    mode) to keep them driving their pre-existing level.
    
    Seeding the shadow reads the PIO registers at probe time, which requires
    the bus clock to be enabled.  The clock was only requested at the very
    end of probe, after devm_pinctrl_register() had already claimed the pin
    hogs described in the device tree - which mux pins, and thus access
    registers, with the clock still gated.  Move the request ahead of both.
    Boards whose bootloader leaves the PIO clock running are unaffected,
    which is why the pre-existing hog problem has gone unnoticed since
    commit 950707c0eb5c ("pinctrl: sunxi: add clock support").
    
    Fixes: df7b34f4c3d2 ("pinctrl: sunxi: Fix gpio_set behaviour")
    Cc: [email protected]
    Signed-off-by: Ilya Titov <[email protected]>
    Signed-off-by: Linus Walleij <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

pinctrl: tegra238: Fix register bank for AON pin groups [+ + +]
Author: Prathamesh Shete <[email protected]>
Date:   Wed Sep 16 06:51:32 2026 +0000

    pinctrl: tegra238: Fix register bank for AON pin groups
    
    [ Upstream commit ae2c5bf969573708cd5b6bb6393631255d83c3fe ]
    
    The AON pin controller has a single register region and therefore,
    the bank defined in the tegra238_functions[] and tegra238_aon_groups[]
    for the AON pin groups must be 0. However, commit 25cac7292d49
    ("pinctrl: tegra: Add Tegra238 pinmux driver") incorrectly specified the
    bank for these pins as 1 and not 0. This means that in the
    tegra_pinctrl_probe() function we use an invalid index when accessing
    the pmx->regs[] array which causes an incorrect address to be used for
    accessing the pinmux registers.
    
    Fix this by correcting the bank for the AON pin groups.
    
    Fixes: 25cac7292d49 ("pinctrl: tegra: Add Tegra238 pinmux driver")
    Signed-off-by: Prathamesh Shete <[email protected]>
    Signed-off-by: Linus Walleij <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
PM: hibernate: Freeze kernel threads after image preallocation [+ + +]
Author: Florian Schmaus <[email protected]>
Date:   Sun Sep 20 16:35:25 2026 +0200

    PM: hibernate: Freeze kernel threads after image preallocation
    
    [ Upstream commit 41112a787f9182c7f2d27122817861e3f22ef928 ]
    
    Commit 783c81098445 ("PM: hibernate: call preallocate_image() after freeze
    prepare") moved hibernate_preallocate_memory() after dpm_prepare() so
    that device drivers have the opportunity to release pinned/unswappable
    memory during their ->prepare() callback before memory is preallocated
    for the snapshot image.
    
    However, that commit also placed hibernate_preallocate_memory() after
    freeze_kernel_threads(). While it was assumed during review that swap
    I/O submitted via submit_bio() is synchronous and would not depend on
    frozen kernel threads, this does not hold in practice. Calling
    hibernate_preallocate_memory() with kernel threads frozen leads to
    intermittent deadlocks during hibernation.
    
    Inside hibernate_preallocate_memory(), shrink_all_memory() is invoked with
    .may_writepage = 1 and .may_swap = 1 to aggressively reclaim and swap out
    pages. Any writeback or swap I/O that relies on freezable kernel threads,
    block device helpers, or WQ_FREEZABLE workqueues (such as those in storage
    drivers, device mapper, or filesystems) deadlocks waiting on tasks that
    are stuck in the refrigerator.
    
    Fix this by reordering hibernation_snapshot():
     1. Call dpm_prepare(PMSG_FREEZE) first, allowing device drivers to
        release pinned resources while kernel threads are still active.
     2. Call hibernate_preallocate_memory() second, performing page
        reclaim and swapout while storage layers, workqueues, and kernel
        threads are alive.
     3. Call freeze_kernel_threads() third, only after all memory
        preallocation and swap I/O have completed.
    
    Additionally, restore the call to swsusp_free() in the cleanup path so
    that preallocated image memory is properly freed if freeze_kernel_threads()
    fails or if TEST_FREEZER is enabled.
    
    Fixes: 783c81098445 ("PM: hibernate: call preallocate_image() after freeze prepare")
    Signed-off-by: Florian Schmaus <[email protected]>
    Reviewed-by: Mario Limonciello (AMD) <[email protected]>
    Tested-by: Matthew Leach <[email protected]>
    Reviewed-by: Matthew Leach <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Rafael J. Wysocki <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
rds: ib: Clear the sg list when mapping an MR fails [+ + +]
Author: Dongliang Qin <[email protected]>
Date:   Tue Sep 22 11:15:45 2026 +0800

    rds: ib: Clear the sg list when mapping an MR fails
    
    commit 58eb1b3325edac42dc6df72c80962bd53a3c8ca7 upstream.
    
    rds_ib_map_frmr() stores the caller's scatterlist in the MR before DMA
    mapping and registration can fail. On failure, __rds_rdma_map() unpins
    the pages and frees the scatterlist, but rds_ib_free_frmr() can still
    return the MR to the pool with the stale pointer set.
    
    This leaves the pool with a dangling scatterlist and can lead to local
    privilege escalation. KASAN detects the resulting use-after-free when the
    MR is later torn down:
    
    BUG: KASAN: slab-use-after-free in __rds_ib_teardown_mr
    Read of size 8
    
    Call Trace:
     __rds_ib_teardown_mr
     rds_ib_unreg_frmr
     rds_ib_flush_mr_pool
     rds_ib_flush_mrs
     rds_free_mr
     rds_setsockopt
    
    Store the scatterlist in the MR only after DMA mapping succeeds. If DMA
    mapping fails, return directly while the MR fields remain clear; the caller
    keeps ownership of the scatterlist and its pinned pages. If a later
    registration step fails, unmap the scatterlist and clear the MR fields
    before returning.
    
    Fixes: 1659185fb4d0 ("RDS: IB: Support Fastreg MR (FRMR) memory registration mode")
    Cc: [email protected]
    Signed-off-by: Dongliang Qin <[email protected]>
    Reviewed-by: Allison Henderson <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
remoteproc: qcom_q6v5_adsp: Fix iommu_unmap() usage [+ + +]
Author: Mostafa Saleh <[email protected]>
Date:   Thu Aug 27 20:30:55 2026 +0000

    remoteproc: qcom_q6v5_adsp: Fix iommu_unmap() usage
    
    [ Upstream commit 0d8e2195bce6f08c1c53c5ef4d7347fe46418101 ]
    
    During adsp_map_carveout, the IOVA is computed by combining the
    physical address and the SID:
        iova = adsp->mem_phys | (sid << 32);
    
    However, adsp_unmap_carveout() uses the physical address and not
    the IOVA in iommu_unmap(), causing the unmap to fail or leak
    mappings because the address doesn't match the original IOVA.
    
    Cache the constructed IOVA within the qcom_adsp device struct
    during mapping and use it during unmapping.
    
    Fixes: f22eedff28af ("remoteproc: qcom: Add support for memory sandbox")
    Signed-off-by: Mostafa Saleh <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Bjorn Andersson <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
Revert "drm/virtio: Allow importing prime buffers when 3D is enabled" [+ + +]
Author: Dmitry Osipenko <[email protected]>
Date:   Fri Sep 11 17:42:03 2026 +0300

    Revert "drm/virtio: Allow importing prime buffers when 3D is enabled"
    
    [ Upstream commit 1e3b08de63274d0b009e99ef51cd6a9c0c6bf08c ]
    
    Guest userspace may import udmabuf to vrend. Vrend doesn't support guest
    blobs, and thus, further 3d operations with the imported blob are failing.
    Typical scenario of the problem shown with mouse cursor RGBA image imported
    into virtio-gpu, which previously was rejected by virtio-gpu driver.
    Revert enabling guest blobs importing into vrend to fix the regression.
    
    Link: https://gitlab.freedesktop.org/virgl/virglrenderer/-/work_items/674
    Fixes: df4dc947c46b ("drm/virtio: Allow importing prime buffers when 3D is enabled")
    Signed-off-by: Dmitry Osipenko <[email protected]>
    Reviewed-by: Val Packett <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

 
RISC-V: KVM: Fix HSM hart status error propagation [+ + +]
Author: Tan Chi <[email protected]>
Date:   Mon Sep 14 11:11:46 2026 +0800

    RISC-V: KVM: Fix HSM hart status error propagation
    
    commit 41e81f7e3ef96594fb840445343c0ee7723aa550 upstream.
    
    kvm_sbi_hsm_vcpu_get_status() returns SBI_ERR_INVALID_PARAM when
    the requested hart does not exist. However, the HART_STATUS case
    returns from the SBI handler without storing this error in
    retdata->err_val.
    
    As a result, a guest querying the status of a non-existent hart
    observes SBI_SUCCESS instead of SBI_ERR_INVALID_PARAM.
    
    Use the common SBI error handling path for HART_STATUS after
    saving a valid hart state in retdata->out_val. This preserves
    the returned error when kvm_sbi_hsm_vcpu_get_status() fails.
    
    Fixes: bae0dfd74e01 ("RISC-V: KVM: Modify SBI extension handler to return SBI error code")
    Cc: [email protected]
    Assisted-by: LLM
    Signed-off-by: Tan Chi <[email protected]>
    Reviewed-by: Anup Patel <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Anup Patel <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

RISC-V: KVM: Fix perf-backed counter accounting across stop and read [+ + +]
Author: SeungJu Cheon <[email protected]>
Date:   Tue Aug 25 17:37:19 2026 +0900

    RISC-V: KVM: Fix perf-backed counter accounting across stop and read
    
    [ Upstream commit c7e2cc38c56142cdab25e6f73602a8222bf9479b ]
    
    pmu_ctr_read() adds the event count returned by perf_event_read_value()
    to counter_val, which can accumulate the same count repeatedly across
    reads. kvm_riscv_vcpu_pmu_ctr_stop() also leaves counter_val stale by
    not folding the current event count into it.
    
    Make reads of perf-backed counters side-effect free, and use
    perf_event_pause() when stopping a counter to fold the current event
    count into counter_val while resetting it. This preserves the counter
    value across stop/start and lets the snapshot path use counter_val
    directly.
    
    Fixes: 0cb74b65d2e5 ("RISC-V: KVM: Implement perf support without sampling")
    Signed-off-by: SeungJu Cheon <[email protected]>
    Reviewed-by: Anup Patel <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Anup Patel <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

RISC-V: KVM: Fix the conversion between vsip and hvip [+ + +]
Author: Yicong Yang <[email protected]>
Date:   Tue Aug 4 21:40:18 2026 +0800

    RISC-V: KVM: Fix the conversion between vsip and hvip
    
    [ Upstream commit 52c6b7d20d3e791a9e75aa2be2a467990154c7cd ]
    
    Per AIA spec 1.0 Section 6.3.2, the interrupt numbers 13-63
    shares same bit position between related VS shadow CSRs and
    hypervisor CSRs. So there's a shift only for SSI, STI and
    SEI interrupt.
    
    Currently the KVM always does a shift for all the interrupts
    (include LCOFI with number 13) when doing the conversion
    between vsip and hvip. Fix this by only doing shift the SSI,
    STI and SEI. Add wrappers for doing the conversion between
    vsip and hvip.
    
    Fixes: 16b0bde9a37c ("RISC-V: KVM: Add perf sampling support for guests")
    Signed-off-by: Yicong Yang <[email protected]>
    Reviewed-by: Anup Patel <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Anup Patel <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

RISC-V: KVM: Preserve firmware counter value across stop/start [+ + +]
Author: SeungJu Cheon <[email protected]>
Date:   Tue Aug 25 17:37:17 2026 +0900

    RISC-V: KVM: Preserve firmware counter value across stop/start
    
    [ Upstream commit 8b3fd1a8b305321171602bfa7c41212441cf69e4 ]
    
    Firmware events accumulate in kvpmu->fw_event[].value while running,
    but counter stop only clears fw_event[].started without saving the
    value back to pmc->counter_val. A subsequent counter start without
    SBI_PMU_START_FLAG_SET_INIT_VALUE reloads the stale counter_val into
    fw_event[].value, losing all events counted so far.
    
    Save fw_event[].value into counter_val when actually stopping a
    running counter, and remove the now redundant synchronization from
    the snapshot path.
    
    Fixes: badc386869e2c ("RISC-V: KVM: Support firmware events")
    Signed-off-by: SeungJu Cheon <[email protected]>
    Reviewed-by: Anup Patel <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Anup Patel <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

RISC-V: KVM: Propagate interrupted G-stage faults [+ + +]
Author: Xie Bo <[email protected]>
Date:   Mon Aug 10 13:15:44 2026 +0800

    RISC-V: KVM: Propagate interrupted G-stage faults
    
    commit f41fb17143df890855d3de980f8eb92dfd595817 upstream.
    
    __kvm_faultin_pfn() reports an interrupted host page fault with
    KVM_PFN_ERR_SIGPENDING. RISC-V currently handles it as a generic
    error PFN and returns -EFAULT.
    
    Return -EINTR for the signal-pending sentinel so callers can distinguish
    an interrupted fault from an invalid userspace mapping. Do not log the
    expected interruption as a vCPU exit error.
    
    Fixes: 9d05c1fee837 ("RISC-V: KVM: Implement stage2 page table programming")
    Cc: [email protected]
    Signed-off-by: Xie Bo <[email protected]>
    Reviewed-by: Anup Patel <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Anup Patel <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

RISC-V: KVM: Release unused page after MMU invalidation [+ + +]
Author: Xie Bo <[email protected]>
Date:   Mon Aug 10 13:15:43 2026 +0800

    RISC-V: KVM: Release unused page after MMU invalidation
    
    commit ed54fdb460a6e81d5f8f38388d1f10b385dd4a58 upstream.
    
    If an MMU invalidation races with a G-stage fault, the fault handler skips
    installing the page but leaves ret set to zero. As a result,
    kvm_release_faultin_page() treats the page as used and can unnecessarily
    mark it dirty.
    
    Track the invalidation retry separately and release the page as unused,
    while preserving the existing return value so that the vCPU retries the
    fault.
    
    Fixes: 2ed90cb0938a ("KVM: RISC-V: Retry fault if vma_lookup() results become invalid")
    Cc: [email protected]
    Signed-off-by: Xie Bo <[email protected]>
    Reviewed-by: Anup Patel <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Anup Patel <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

RISC-V: KVM: Report snapshot write failure to the guest [+ + +]
Author: SeungJu Cheon <[email protected]>
Date:   Tue Aug 25 17:37:18 2026 +0900

    RISC-V: KVM: Report snapshot write failure to the guest
    
    [ Upstream commit 057dd2639ceae79adced5d8fe52c32d562edcb3a ]
    
    If kvm_vcpu_write_guest() fails while updating the PMU snapshot area
    on counter stop, the guest may receive SBI_SUCCESS without the
    snapshot being updated, leaving stale data in shared memory.
    
    Return SBI_ERR_FAILURE when the snapshot write fails.
    
    Fixes: c2f41ddbcdd7 ("RISC-V: KVM: Implement SBI PMU Snapshot feature")
    Signed-off-by: SeungJu Cheon <[email protected]>
    Reviewed-by: Anup Patel <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Anup Patel <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

RISC-V: KVM: Serialize IMSIC attributes with vCPU migration [+ + +]
Author: Xie Bo <[email protected]>
Date:   Mon Aug 10 09:21:15 2026 +0800

    RISC-V: KVM: Serialize IMSIC attributes with vCPU migration
    
    commit 8ae12ccaec6ec74945d8c1ef39f2c1b8df779abc upstream.
    
    KVM device ioctls are not serialized against KVM_RUN. As a result,
    kvm_riscv_aia_imsic_rw_attr() can snapshot the physical CPU and HGEI of
    an IMSIC VS-file before a concurrent vCPU migration releases it.
    
    The HGEI can then be allocated to another vCPU before imsic_vsfile_rw()
    uses the stale tuple. A GET or SET attribute may consequently access the
    new owner's interrupt file.
    
    Serialize the entire IMSIC attribute operation with the target vCPU
    mutex. This prevents the VS-file from being migrated and recycled until
    the attribute access completes. Acquire the mutex killably so that the
    device ioctl remains interruptible while waiting for KVM_RUN to finish.
    
    Fixes: db8b7e97d613 ("RISC-V: KVM: Add in-kernel virtualization of AIA IMSIC")
    Cc: [email protected]
    Signed-off-by: Xie Bo <[email protected]>
    Reviewed-by: Anup Patel <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Anup Patel <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

RISC-V: KVM: Synchronize hrtimer callback during teardown [+ + +]
Author: Myeonghun Pak <[email protected]>
Date:   Sat Aug 1 01:35:50 2026 +0900

    RISC-V: KVM: Synchronize hrtimer callback during teardown
    
    commit aaad136d56d91252517272b68cd533e5714698d5 upstream.
    
    The non-Sstc hrtimer callback clears next_set before its final uses of
    the enclosing vCPU.  If teardown observes next_set as false while the
    callback is still running, kvm_riscv_vcpu_timer_cancel() skips
    hrtimer_cancel() and kvm_destroy_vcpus() can free the vCPU before the
    callback enters kvm_riscv_vcpu_set_interrupt().
    
    A guest can arm the timer with SBI TIME and request shutdown with SBI
    legacy shutdown or SRST.  A VMM that honors KVM_EXIT_SYSTEM_EVENT and
    destroys the VM supplies the teardown side of the race; no post-launch
    host ioctl is needed to arm or request teardown.
    
    On upstream master 62cc90241548, generic KASAN reported:
    
      BUG: KASAN: slab-use-after-free in do_raw_spin_lock
      Write of size 4 at addr ff60000005e58898
    
      kvm_riscv_vcpu_set_interrupt
      kvm_riscv_vcpu_hrtimer_expired
      __hrtimer_run_queues
      hrtimer_interrupt
    
    The object was allocated by KVM_CREATE_VCPU and freed concurrently by:
    
      kvm_destroy_vcpus
      kvm_arch_destroy_vm
      kvm_destroy_vm
      __fput
    
    For deterministic validation, I added mdelay(1000) immediately after
    the existing next_set = false assignment.  This only widens the
    existing post-clear callback window.  A no-delay trace build naturally
    reached the callback-after-teardown-start/before-deinit ordering in 12
    of 200 runs, but 1,500 stock-kernel stress iterations did not produce a
    KASAN report, so natural reproduction is timing-sensitive.
    
    Always invoke hrtimer_cancel() for an initialized timer.  Preserve the
    existing -EINVAL result when the timer is no longer set, but only after
    synchronizing with a running callback.
    
    With this patch, hrtimer_cancel() blocked for the full widened callback
    window before vCPU destruction.  KASAN reported no error in 100
    fixed-and-widened runs or 200 fix-only timing-sweep runs.
    
    Fixes: 3a9f66cb25e1 ("RISC-V: KVM: Add timer functionality")
    Cc: [email protected]
    Assisted-by: OpenAI:GPT-5.6
    Signed-off-by: Myeonghun Pak <[email protected]>
    Reviewed-by: Anup Patel <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Anup Patel <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
s390/cio: Check pmcw.dnv before pmcw.ena in I/O entry points [+ + +]
Author: Vineeth Vijayan <[email protected]>
Date:   Thu Sep 10 11:32:02 2026 +0200

    s390/cio: Check pmcw.dnv before pmcw.ena in I/O entry points
    
    [ Upstream commit f6f2985eabdb2bfdc82ce90a1ea3ec53ba795f34 ]
    
    The device number valid (dnv) bit in the PMCW must be checked before
    acting on any other PMCW fields for IO-type subchannels. A subchannel
    with dnv=0 has no valid device number associated, making it meaningless
    to evaluate the enabled (ena) state or issue any I/O instruction against
    it.
    
    Reported-by: William Bezenah <[email protected]>
    Signed-off-by: Vineeth Vijayan <[email protected]>
    Reviewed-by: Peter Oberparleiter <[email protected]>
    Fixes: 8c58a229688c ("s390/cio: Do not unregister the subchannel based on DNV")
    Signed-off-by: Heiko Carstens <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

s390/cio: Fix cio_update_schib() to not cache invalid schib [+ + +]
Author: Vineeth Vijayan <[email protected]>
Date:   Thu Sep 10 11:32:01 2026 +0200

    s390/cio: Fix cio_update_schib() to not cache invalid schib
    
    [ Upstream commit 29d9e5835d89223aa913dcf7b942cc1c148bdd25 ]
    
    When pmcw.dnv is 0, the contents of all SCHIB fields are unpredictable.
    Zero sch->schib in that case to prevent subsequent code from making
    decisions based on unpredictable data.
    
    Reported-by: William Bezenah <[email protected]>
    Signed-off-by: Vineeth Vijayan <[email protected]>
    Reviewed-by: Peter Oberparleiter <[email protected]>
    Fixes: 8c58a229688c ("s390/cio: Do not unregister the subchannel based on DNV")
    Signed-off-by: Heiko Carstens <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

s390/cio: Fix NULL pointer dereference in ccw_device_get_util_str() [+ + +]
Author: Vineeth Vijayan <[email protected]>
Date:   Tue Sep 22 22:48:39 2026 +0200

    s390/cio: Fix NULL pointer dereference in ccw_device_get_util_str()
    
    commit 5b76268dac968612f7283d59b539036de955b7d9 upstream.
    
    The channel path registry entry associated with a CHPID may be removed
    while the subchannel's PMCW still references that CHPID. In this case,
    chpid_to_chp() can return NULL, leading to a NULL pointer dereference.
    
    Add the missing NULL check before dereferencing the returned pointer.
    
    Fixes: 199652309a4d ("s390/cio: add helper to query utility strings per given ccw device")
    Cc: [email protected]
    Signed-off-by: Vineeth Vijayan <[email protected]>
    Reviewed-by: Peter Oberparleiter <[email protected]>
    Signed-off-by: Heiko Carstens <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

s390/cio: Guard PMCW field accesses with dnv check [+ + +]
Author: Vineeth Vijayan <[email protected]>
Date:   Thu Sep 10 11:32:03 2026 +0200

    s390/cio: Guard PMCW field accesses with dnv check
    
    [ Upstream commit 9590f4d83880dfb5a81906e48e72779248fbe8f0 ]
    
    When PMCW.DNV is 0, no I/O device is associated with the subchannel.
    However, several code paths access PMCW fields directly from the cached
    sch->schib without first invoking the update helper. Add explicit DNV
    validation before accessing PMCW fields from the cached SCHIB to avoid
    using invalid data.
    
    Reported-by: William Bezenah <[email protected]>
    Signed-off-by: Vineeth Vijayan <[email protected]>
    Reviewed-by: Peter Oberparleiter <[email protected]>
    Fixes: 8c58a229688c ("s390/cio: Do not unregister the subchannel based on DNV")
    Signed-off-by: Heiko Carstens <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
s390/cmf: Fix virtual vs physical address confusion [+ + +]
Author: Peter Oberparleiter <[email protected]>
Date:   Mon Sep 21 15:16:48 2026 +0200

    s390/cmf: Fix virtual vs physical address confusion
    
    commit f4d04425e66af2ecee9d1a49ae0484436f3c2fd1 upstream.
    
    The measurement block address is an absolute address. Define the
    associated schib_config and schib fields as dma64_t to enable automatic
    detection of incorrect assignments. Also add the missing virt_to_dma64()
    translation.
    
    Without this fix, a wrong address will be used by firmware when storing
    extended format channel measurement data on kernels built with
    CONFIG_RANDOMIZE_IDENTITY_BASE=y.
    
    Fixes: 14edd0d73bfe ("s390/cmf: fix virtual vs physical address confusion")
    Cc: [email protected]
    Signed-off-by: Peter Oberparleiter <[email protected]>
    Reviewed-by: Heiko Carstens <[email protected]>
    Signed-off-by: Heiko Carstens <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
s390/debug: Do not register views for failed static debug areas [+ + +]
Author: Mikhail Zaslonko <[email protected]>
Date:   Fri Sep 18 17:13:51 2026 +0200

    s390/debug: Do not register views for failed static debug areas
    
    [ Upstream commit 28e29992b034acffc9342df216c06097825ce610 ]
    
    __REGISTER_STATIC_DEBUG_INFO() calls debug_register_view()
    unconditionally, even when debug_register_static() has failed. In that
    case _debug_register() was never reached and id->debugfs_root_entry is
    still NULL, so debugfs_create_file() places the view file in the debugfs
    root directory. For sclp_err this leaves a /sys/kernel/debug/hex_ascii
    file with nothing to indicate which debug log it belongs to.
    
    debug_register_static() is not exported and the macro is its only
    caller, so let it return an error code and skip the view registration
    when it fails. No debugfs files are created for such an area then.
    
    Reproduce by booting with s390dbf=sclp_err::100000000. The sclp_err
    registration fails, no s390dbf/sclp_err/ directory is created, and a
    hex_ascii file appears in the debugfs root instead.
    
    Fixes: d72541f94512 ("s390/debug: add early tracing support")
    Signed-off-by: Mikhail Zaslonko <[email protected]>
    Reviewed-by: Heiko Carstens <[email protected]>
    Signed-off-by: Heiko Carstens <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

s390/debug: Fix NULL pointer dereference in debug_info_copy() [+ + +]
Author: Mikhail Zaslonko <[email protected]>
Date:   Wed Sep 16 18:06:26 2026 +0200

    s390/debug: Fix NULL pointer dereference in debug_info_copy()
    
    [ Upstream commit 012bfcd5a51082d5a65f096dfb9ca652b5267965 ]
    
    When debug_register_static() fails, it clears areas, active_pages and
    active_entries but leaves the area bounds unchanged. Copying such an
    area, either by opening its view file or via debug_dump(), makes
    debug_info_copy() dereference the NULL pointers.
    
    Skip the copy loop when the source has no areas.
    
    Closes: https://lore.kernel.org/r/[email protected]
    Fixes: d72541f94512 ("s390/debug: add early tracing support")
    Signed-off-by: Mikhail Zaslonko <[email protected]>
    Reviewed-by: Peter Oberparleiter <[email protected]>
    Acked-by: Heiko Carstens <[email protected]>
    Signed-off-by: Heiko Carstens <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
s390/pci/docs: Fix sriov_numvfs attribute name [+ + +]
Author: Karl Mehltretter <[email protected]>
Date:   Mon Sep 7 07:58:47 2026 +0200

    s390/pci/docs: Fix sriov_numvfs attribute name
    
    [ Upstream commit 4525a911049543c23885a540a788d13be318a486 ]
    
    The attribute is sriov_numvfs (drivers/pci/iov.c); the document names it
    sriov_numvf, which does not exist.
    
    Use sriov_numvfs.
    
    Fixes: de267a7c71ba ("s390/pci: Documentation for zPCI")
    Assisted-by: LLM
    Signed-off-by: Karl Mehltretter <[email protected]>
    Reviewed-by: Randy Dunlap <[email protected]>
    Signed-off-by: Heiko Carstens <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
s390/pci: Don't report recovery success on skipped recovery [+ + +]
Author: Niklas Schnelle <[email protected]>
Date:   Wed Sep 16 17:14:13 2026 +0200

    s390/pci: Don't report recovery success on skipped recovery
    
    commit 3a43be7a1fd06a35cf9e621b88283b3b6e7d281c upstream.
    
    When a PCI device is already in the permanent failure state, recovery is
    skipped, but the SCLP recovery report still shows success. Fix this by
    changing the status string to explicitly state that recovery was skipped
    due to permanent failure.
    
    Cc: [email protected]
    Fixes: 4ec6054e7321 ("s390/pci: Report PCI error recovery results via SCLP")
    Signed-off-by: Niklas Schnelle <[email protected]>
    Reviewed-by: Benjamin Block <[email protected]>
    Reviewed-by: Farhan Ali <[email protected]>
    Signed-off-by: Heiko Carstens <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

s390/pci: Fix leak of struct pci_dev reference in zpci_report_status() [+ + +]
Author: Niklas Schnelle <[email protected]>
Date:   Wed Sep 16 17:14:10 2026 +0200

    s390/pci: Fix leak of struct pci_dev reference in zpci_report_status()
    
    commit 09b7040a1b79f4f61cdad6d9af972e045bb7c498 upstream.
    
    In zpci_report_status(), a reference to the pdev associated with the
    zdev being reported about is acquired using pci_get_slot(). This
    reference needs to be dropped with pci_dev_put(), but this call is
    missing, thus leaking the reference. On subsequent hot unplug, this will
    cause the struct pci_dev to not be released, leaking memory and
    preventing reattach.
    
    At the same time, the only existing caller already holds a pdev
    reference. So instead of reacquiring and then dropping another reference,
    simply pass the existing pdev pointer to zpci_report_status(). This gets
    rid of the need for pci_get_slot() as well as the zdev->zbus check.
    
    Cc: [email protected]
    Fixes: 4ec6054e7321 ("s390/pci: Report PCI error recovery results via SCLP")
    Signed-off-by: Niklas Schnelle <[email protected]>
    Reviewed-by: Benjamin Block <[email protected]>
    Reviewed-by: Farhan Ali <[email protected]>
    Signed-off-by: Heiko Carstens <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

s390/pci: Fix missing device lock in zpci_report_status() [+ + +]
Author: Niklas Schnelle <[email protected]>
Date:   Wed Sep 16 17:14:11 2026 +0200

    s390/pci: Fix missing device lock in zpci_report_status()
    
    commit 0261aef4b15efcee2860ab857e5cb05e9bfa47b0 upstream.
    
    When pdev is non-NULL, zpci_report_status() accesses the device's driver.
    To get a consistent state matching the recovery, the device lock needs to
    be held. Do so by expanding the existing device lock critical section.
    
    The lock only needs to be held when the pdev is non-NULL, so extract
    the pdev-specific reporting into a helper function which also adds a
    lockdep assertion to detect calls without the device lock held.
    
    Cc: [email protected]
    Fixes: 4ec6054e7321 ("s390/pci: Report PCI error recovery results via SCLP")
    Signed-off-by: Niklas Schnelle <[email protected]>
    Reviewed-by: Benjamin Block <[email protected]>
    Reviewed-by: Farhan Ali <[email protected]>
    Signed-off-by: Heiko Carstens <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

s390/pci: Report SCLP status on error events when no pdev is associated [+ + +]
Author: Niklas Schnelle <[email protected]>
Date:   Wed Sep 16 17:14:12 2026 +0200

    s390/pci: Report SCLP status on error events when no pdev is associated
    
    commit a1120bea9bc8ea9d9ab2f9904a63b9a228bf2ffc upstream.
    
    With commit 4ec6054e7321 ("s390/pci: Report PCI error recovery results
    via SCLP") SCLP reports are generated when recovery is performed in
    response to an error event. If such an error event arrives but no pdev
    is currently associated with the zdev, e.g. because it was removed or
    not yet probed, no report is generated. Fix this by generating a report
    specific to an error event for a zdev without an associated pdev.
    
    Cc: [email protected]
    Fixes: 4ec6054e7321 ("s390/pci: Report PCI error recovery results via SCLP")
    Signed-off-by: Niklas Schnelle <[email protected]>
    Reviewed-by: Benjamin Block <[email protected]>
    Reviewed-by: Farhan Ali <[email protected]>
    Signed-off-by: Heiko Carstens <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
s390/uv: Fix loop condition in uv_find_secrets [+ + +]
Author: Steffen Eiden <[email protected]>
Date:   Wed Aug 12 17:55:19 2026 +0200

    s390/uv: Fix loop condition in uv_find_secrets
    
    [ Upstream commit d12ce6bce5ec5175c3581e01c71e7a5abb286d9b ]
    
    Systems with more than 85 UV secrets got -ENOENT for any secret past the
    first page.
    
    Fix this by setting the start index at the beginning of the loop in
    uv_find_secret() and not at the end. First test if there are more
    secrets left by comparing start_idx with list->next_secret_idx, and then
    set the start index to the next secret index.
    
    Fixes: 7c9137af2042 ("s390/uv: Retrieve UV secrets support")
    Acked-by: Claudio Imbrenda <[email protected]>
    Reviewed-by: Christoph Schlameuss <[email protected]>
    Signed-off-by: Steffen Eiden <[email protected]>
    Signed-off-by: Claudio Imbrenda <[email protected]>
    Message-ID: <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

s390/uv: Prevent potential out-of-bounds read [+ + +]
Author: Steffen Eiden <[email protected]>
Date:   Wed Aug 12 17:55:20 2026 +0200

    s390/uv: Prevent potential out-of-bounds read
    
    [ Upstream commit f47190b08b71e8482072978373ee88cb2dfbdaf4 ]
    
    When the system has more than 85 secrets, the uv_secret_list struct
    array only holds up to 85 items per page, resulting in an out of bounds
    read in find_secret_in_page if the targeted secret is in the next page
    or not stored at all.
    
    Fix this by looping over the number of stored secrets which is the
    per sub-list count of stored secrets and not the overall count.
    
    Fixes: 7c9137af2042 ("s390/uv: Retrieve UV secrets support")
    Signed-off-by: Steffen Eiden <[email protected]>
    Reviewed-by: Christoph Schlameuss <[email protected]>
    Signed-off-by: Claudio Imbrenda <[email protected]>
    Message-ID: <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
s390/vfio-ap: fix KVM GISC and page leak when queue removed from host config [+ + +]
Author: Anthony Krowiak <[email protected]>
Date:   Tue Aug 18 15:33:49 2026 -0400

    s390/vfio-ap: fix KVM GISC and page leak when queue removed from host config
    
    commit 65e05ec252a9b79e75930d3c4dd42d8877db04c5 upstream.
    
    Three related problems exist in the handling of KVM interrupt and page
    resources when a queue is removed from the host's AP configuration
    while assigned to a mediated device (mdev).
    
    Problem 1:
    ~~~~~~~~~
    AP_RESPONSE_Q_NOT_AVAIL not handled in vfio_ap_mdev_reset_queue()
    
    When the AP bus removes a queue device whose adapter or domain has
    been removed from the host's AP configuration,
    vfio_ap_mdev_remove_queue() is called. If the queue is still in the
    host's AP configuration at that point, it calls
    vfio_ap_mdev_reset_queue(), which issues a PQAP(ZAPQ). Since the
    adapter is already gone from the host configuration, ap_zapq() returns
    AP_RESPONSE_Q_NOT_AVAIL (0x01). This response code is not handled in
    vfio_ap_mdev_reset_queue()'s switch statement and falls through to
    the default case, which issues a WARN but does not call
    vfio_ap_free_aqic_resources(). As a result, if IRQ handling was
    enabled for the queue by the guest, the KVM GISC registration and
    the pinned guest page holding the notification indicator byte (NIB)
    are both leaked.
    
    This is fixed by adding AP_RESPONSE_Q_NOT_AVAIL to the same case as
    AP_RESPONSE_DECONFIGURED and AP_RESPONSE_CHECKSTOPPED in
    vfio_ap_mdev_reset_queue(). Like those response codes, Q_NOT_AVAIL
    indicates the queue is not operational and no further reset attempts
    are possible; the correct action is to free the IRQ resources
    immediately.
    
    Problem 2:
    ~~~~~~~~~
     AP_RESPONSE_Q_NOT_AVAIL not handled in apq_status_check()
    
    In vfio_ap_mdev_reset_queue(), there are four cases that indicate a queue
    reset has not yet completed, in which case apq_reset_check() is queued to
    a work queue to verify completion of the reset operation. This function
    uses the PQAP(TAPQ) function to get the queue's status and calls
    apq_status_check() to verify whether the reset has completed, failed or
    needs to be executed again. As described in Problem #1 above,
    apq_reset_check() does not specifically check for AP_RESPONSE_Q_NOT_AVAIL,
    thereby potentially leaking KVM GISC registration and the pinned guest page
    holding the NIB.
    
    This is fixed by adding a case statement for AP_RESPONSE_Q_NOT_AVAIL to
    apq_status_check() and returning -ENODEV for that case. The caller,
    apq_reset_check() will then check for this return code and call
    vfio_ap_free_aqic_resources() to prevent the leak.
    
    Problem 3:
    ~~~~~~~~~
    vfio_ap_free_aqic_resources() leaks saved_isc when kvm is NULL
    
    vfio_ap_free_aqic_resources() guards the call to
    kvm_s390_gisc_unregister() with:
    
        if (q->saved_isc != VFIO_AP_ISC_INVALID &&
            !WARN_ON(!(q->matrix_mdev && q->matrix_mdev->kvm)))
    
    If matrix_mdev->kvm is NULL -- which can happen when
    vfio_ap_mdev_unset_kvm() has already run and cleared kvm before a
    subsequent cleanup path reaches this function -- the WARN_ON fires
    and the entire block is skipped. This leaves q->saved_isc set to a
    non-invalid value, creating a potential double-free on any subsequent
    call to this function.
    
    When kvm is NULL the KVM guest is already torn down, so
    kvm_s390_gisc_unregister() need not and cannot be called; however,
    q->saved_isc must always be cleared. Fix this by separating the
    kvm_s390_gisc_unregister() call from the q->saved_isc reset. The
    WARN_ON now guards only the genuinely impossible case of matrix_mdev
    being NULL. A NULL kvm is handled gracefully by skipping only the
    unregister call, and q->saved_isc = VFIO_AP_ISC_INVALID is set
    unconditionally whenever saved_isc was not already invalid.
    
    Additionally, add an else clause to the host-config check in
    vfio_ap_mdev_remove_queue() to call vfio_ap_free_aqic_resources()
    directly when the queue is not in the host's AP configuration. This
    serves as a backstop: when the AP bus fires the driver .remove
    callback after an adapter is removed from the host config, the queue
    is by definition no longer addressable, so vfio_ap_mdev_reset_queue()
    would always return Q_NOT_AVAIL. The else clause handles this case
    directly without the unnecessary ap_zapq() call, and ensures cleanup
    occurs even if kvm has already been set to NULL by a prior call to
    vfio_ap_mdev_unset_kvm().
    
    Fixes: b9bd10c43456d ("s390/vfio-ap: do not reset queue removed from host config")
    Cc: [email protected]
    Signed-off-by: Anthony Krowiak <[email protected]>
    Reviewed-by: Matthew Rosato <[email protected]>
    Acked-by: Halil Pasic <[email protected]>
    Signed-off-by: Claudio Imbrenda <[email protected]>
    Message-ID: <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
sched/cache: Decouple sched_cache_group from mm to fix UAF [+ + +]
Author: Tim Chen <[email protected]>
Date:   Mon Sep 21 17:37:24 2026 -0700

    sched/cache: Decouple sched_cache_group from mm to fix UAF
    
    commit 28f9c0e0a0b94c5d3e1b634db545f6e1f94858c5 upstream.
    
    Currently the sched cache grouping is by mm and the scheduling statistics
    sched_cache_stat lives in the mm structure.  This ties the life cycle
    of scheduling stats with mm.
    
    In account_mm_sched(), the scheduling stats are accessed by
    task->mm->sc_stat.  However, a task may be switching mm on one CPU when
    another CPU is running account_mm_sched(), and possibly accessing the
    old mm that was freed.  This problem was found when running tests with
    KASAN by Hyunwoo:
    
      https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
    
    Instead of serializing the mm access by introducing extra acquisition of
    rq lock in the mm free path, extract sched_cache_stat from mm_struct,
    rename it as sched_cache_group and manage its life cycle apart from
    mm_struct with its own ref counting.  This allows us in the next patch
    access sched_cache_group directly from task, and add a refcount
    on sched_cache_group when a task links to it. This prevents the use
    after free issue when accessing stale and released old mm and its
    sched cache stat a task switches to a new mm while account_mm_sched()
    is done elsewhere.
    
    The other benefit of this restructure is in the future, the grouping of
    tasks to a LLC would have the flexibility to be associated with a user
    defined grouping, or cgroup, cookie group, numa_group or others instead
    of just with a single mm address space.
    
    Rename sched_cache_stat to sched_cache_group and turn it into a refcounted
    object allocated from mm_struct.  The mm_struct now holds a pointer
    (sched_cache_grp) to this object instead of embedding it.
    
    Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware load balancing")
    Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/
    Closes: https://lore.kernel.org/all/[email protected]/
    Reported-by: Hyunwoo Kim <[email protected]>
    Reported-by: Zenghui Yu (Huawei) <[email protected]>
    Co-developed-by: Chen Yu <[email protected]>
    Signed-off-by: Chen Yu <[email protected]>
    Signed-off-by: Tim Chen <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Cc: <[email protected]> #7.2.x
    Link: https://patch.msgid.link/91fd1e3266707c865bc9abecfb3e17bc676712df.1790035273.git.tim.c.chen@linux.intel.com
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug [+ + +]
Author: Davi Chaves Azevedo <[email protected]>
Date:   Mon Sep 21 17:37:27 2026 -0700

    sched/cache: Refresh LLC capacity across CPU hotplug, to fix capacity underestimation bug
    
    commit 3cb0243767fd033bdce95f4f1b5882172a2f8119 upstream.
    
    The scheduler scales LLC capacity by the fraction of cache-sharing CPUs
    covered by a domain:
    
      llc_bytes = cache_size * span_weight / shared_weight
    
    During CPU teardown, sched_cpu_deactivate() rebuilds scheduler domains
    before cacheinfo_cpu_pre_down() removes the CPU from shared_cpu_map. The
    new domains therefore use the old sharing weight. The later call to
    sched_update_llc_bytes() looks up the departing CPU's sd_llc, which has
    already been detached, and returns without correcting the surviving CPUs.
    
    On a Ryzen 5 7535U with twelve logical CPUs sharing a 16 MiB LLC,
    offlining one SMT sibling left the remaining CPUs with:
    
      llc_bytes = floor(16777216 * 11 / 12) = 15379114 bytes
    
    The correct capacity is still 16777216 bytes. On systems with active
    cache-aware scheduling, an underestimated capacity can cause
    exceed_llc_capacity() to reject aggregation for a process whose footprint
    would fit. Unchanged cpuset partitions sharing the physical cache can
    also retain stale capacity when a CPU comes online in another partition.
    
    Pass the cache-sharing mask already retained by cacheinfo to the
    scheduler update. Refresh every surviving CPU using its own LLC domain
    so that each partition receives the correct share. This also preserves
    the correction needed as cache-sharing maps grow during boot.
    
    Keep the existing CPU-hotplug and scheduler-domain synchronization. The
    update remains on the hotplug path; no steady-state scheduling operation
    or persistent allocation is added.
    
    Fixes: 7030513a0877 ("sched/cache: Calculate the LLC size and store it in sched_domain")
    Signed-off-by: Davi Chaves Azevedo <[email protected]>
    Signed-off-by: Tim Chen <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Reviewed-by: Chen Yu <[email protected]>
    Reviewed-by: Tim Chen <[email protected]>
    Reviewed-by: K Prateek Nayak <[email protected]>
    Tested-by: Chen Yu <[email protected]>
    Tested-by: K Prateek Nayak <[email protected]>
    Cc: <[email protected]> # v7.2.x
    Link: https://patch.msgid.link/6751d93e15889e624796c74db0bfe66603d60b1b.1790035273.git.tim.c.chen@linux.intel.com
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
sched/core: Account PSI IRQ time to the execution context, not the scheduling context [+ + +]
Author: Zhan Xusheng <[email protected]>
Date:   Fri Sep 18 21:29:15 2026 +0800

    sched/core: Account PSI IRQ time to the execution context, not the scheduling context
    
    [ Upstream commit a0bb6fac53fa7cf1cadb487b43d4c9276a6b82e3 ]
    
    psi_account_irqtime() has two callers which share rq->psi_irq_time, and
    they disagree about the context: __schedule() passes the outgoing rq->curr,
    sched_tick() passes rq->donor.  Under proxy execution the donor is blocked
    on a mutex while rq->curr burns the CPU.
    
    The tick charges PSI_IRQ_FULL to the donor's cgroup and advances the
    timestamp, so the call from __schedule() then finds delta <= 0 and charges
    nothing.  The delta is not counted twice, it lands on the wrong cgroup.
    
    Pass rq->curr, which is what the call read before commit af0c8b2bf67b
    ("sched: Split scheduler and execution contexts") renamed 'curr' to
    'donor' across sched_tick().  Without CONFIG_SCHED_PROXY_EXEC the two rq
    members are a union, so this only changes anything where that option is set,
    and it depends on EXPERT.
    
    Fixes: af0c8b2bf67b ("sched: Split scheduler and execution contexts")
    Signed-off-by: Zhan Xusheng <[email protected]>
    Signed-off-by: Peter Zijlstra (Intel) <[email protected]>
    Signed-off-by: Ingo Molnar <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Sasha Levin <[email protected]>

 
sched_ext: Derive SCX_RQ_IN_WAKEUP from the core enqueue flags [+ + +]
Author: Tejun Heo <[email protected]>
Date:   Wed Sep 16 12:00:24 2026 -1000

    sched_ext: Derive SCX_RQ_IN_WAKEUP from the core enqueue flags
    
    commit df5cdc2c832ca4e8a6d774596b9005558761a403 upstream.
    
    schedule_deferred_locked() skips scheduling a deferred action while
    SCX_RQ_IN_WAKEUP is set and relies on the task_woken_scx() call that follows
    a wakeup enqueue to run it. enqueue_task_scx() sets the flag from the merged
    enqueue flags, which include the flags stashed for a remote activation.
    move_remote_task_to_local_dsq() thus sets SCX_RQ_IN_WAKEUP on the
    destination rq when the moved task was woken up, although no
    task_woken_scx() follows that activation.
    
    An IMMED insert into a busy destination requests a local reenqueue during
    that enqueue. The request gets linked but not scheduled and stays pending
    until an unrelated wakeup or preemption on that CPU runs the deferred
    actions. The IMMED task sits behind the running task in the meantime. If
    nothing runs them before the scheduler is disabled, the request outlives the
    scheduler and points into its freed per-cpu area, which the next scheduler
    dereferences from run_deferred().
    
    Test the core enqueue flags for the wakeup bit. Only the core's wakeup path
    is followed by task_woken_scx().
    
    Fixes: 57ccf5ccdc56 ("sched_ext: Fix enqueue_task_scx() truncation of upper enqueue flags")
    Cc: [email protected] # v7.1+
    Reported-by: Andrea Righi <[email protected]>
    Link: https://lore.kernel.org/all/[email protected]/
    Signed-off-by: Tejun Heo <[email protected]>
    Reviewed-by: Andrea Righi <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
scsi: block: Fix zones_cond out-of-bounds write on zone report [+ + +]
Author: ZHOU Jiaxiang <[email protected]>
Date:   Wed Sep 16 21:58:21 2026 +0800

    scsi: block: Fix zones_cond out-of-bounds write on zone report
    
    [ Upstream commit 7c431d61b69a3fd0784c20aa4cd0b8fb501b5653 ]
    
    blk_revalidate_disk_zones() sizes the zones_cond array from the disk
    capacity and zone size, but the index used by blk_revalidate_zone_cond()
    comes from the device-driven report_zones() walk and is never checked
    against the array size. A device reporting more zones than fit the array
    makes blk_zone_set_cond() write out of bounds.
    
    One way to reach this is a zone count exceeding 32 bits: both
    blk_revalidate_zone_args.nr_zones and struct zoned_disk_info.nr_zones are
    unsigned int, so a disk advertising more than UINT_MAX zones (e.g.  2^32 +
    1024 zones of one 512-byte logical block) gets its zone count truncated to
    a small value, undersizing the array while the report walk keeps counting
    upward.
    
    Check the index against the array size before storing the zone condition,
    and refuse to revalidate when the zone count does not fit 32 bits.
    
    Fixes: 6e945ffb6555 ("block: use zone condition to determine conventional zones")
    Signed-off-by: ZHOU Jiaxiang <[email protected]>
    Reviewed-by: Damien Le Moal <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Martin K. Petersen (Oracle) <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

scsi: libiscsi_tcp: Check the data direction of a Data-In PDU [+ + +]
Author: Yehyeong Lee <[email protected]>
Date:   Sat Aug 1 22:36:35 2026 +0900

    scsi: libiscsi_tcp: Check the data direction of a Data-In PDU
    
    commit bce07e2f37b5e4a427d36fd6b1c14067b27591db upstream.
    
    The Data-In branch of iscsi_tcp_hdr_dissect() resolves the ITT to a task
    and copies the PDU's data segment into that command's scatterlist without
    asking whether the command was reading. iscsi_tcp_r2t_rsp() in the same
    file does ask, and rejects an R2T for a command that is not DMA_TO_DEVICE.
    
    A target that answers a WRITE command's ITT with a Data-In therefore has
    the initiator write target-supplied bytes into the pages that write was
    about to send. Those are the caller's own pinned pages for an O_DIRECT
    write, and page cache pages for a buffered one.
    
    Observed against a test target that emits one 512-byte Data-In naming a 128
    KB write's ITT, after the R2T for that write. With O_DIRECT the caller's
    buffer ends up holding 512 bytes of the target's data while pwrite()
    returns 131072. Buffered is quieter: pwrite() and fsync() both succeed,
    nothing is logged, and reading those blocks back returns the target's bytes
    out of the page cache without a command going on the wire.
    
    Check the direction before using the scatterlist, the way the R2T path
    already does.
    
    Cc: [email protected]
    Signed-off-by: Yehyeong Lee <[email protected]>
    Reviewed-by: Mike Christie <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Fixes: a081c13e39b5 ("[SCSI] iscsi_tcp: split module into lib and lld")
    Signed-off-by: Martin K. Petersen (Oracle) <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

scsi: megaraid_sas: Protect megasas_get_ctrl_info() in megasas_resume() [+ + +]
Author: Bart Van Assche <[email protected]>
Date:   Mon Aug 31 12:27:20 2026 -0700

    scsi: megaraid_sas: Protect megasas_get_ctrl_info() in megasas_resume()
    
    [ Upstream commit 42d1221d321e55afc7bba9109a77aaf5a817c8a3 ]
    
    Protect the megasas_get_ctrl_info() call in megasas_resume() with
    instance->reset_mutex using scoped_guard().
    
    megasas_get_ctrl_info() may release and reacquire instance->reset_mutex.
    Hence, calling this function without holding instance->reset_mutex is not
    safe.
    
    Fixes: c3b10a55abc9 ("scsi: megaraid_sas: Update controller info during resume")
    Cc: Kashyap Desai <[email protected]>
    Cc: Sumit Saxena <[email protected]>
    Cc: Shivasharan S <[email protected]>
    Cc: Chandrakanth patil <[email protected]>
    Signed-off-by: Bart Van Assche <[email protected]>
    Link: https://patch.msgid.link/f06b5ee432b21cf293f0663e15b64f75a84b9fd5.1788204406.git.bvanassche@acm.org
    Signed-off-by: Martin K. Petersen (Oracle) <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

scsi: sd_zbc: Reject disks with too many zones [+ + +]
Author: ZHOU Jiaxiang <[email protected]>
Date:   Wed Sep 16 21:58:22 2026 +0800

    scsi: sd_zbc: Reject disks with too many zones
    
    [ Upstream commit b6ec0f79745967c751c85df373062c8d15e45fc4 ]
    
    sd_zbc_read_zones() computes the number of zones with 64-bit arithmetic and
    stores the result in the unsigned int nr_zones field of struct
    zoned_disk_info, silently truncating counts that exceed 32 bits. The
    truncated count is later used to size per-zone resources, while the device
    may still report more zones than fit.
    
    Moreover, sd_zbc_report_zones() counts the reported zones with a signed int
    zone_idx, which overflows past INT_MAX. Reject devices reporting more than
    INT_MAX zones at scan time; such a device is not realistic for any medium
    that exists today, and accepting it produces inconsistent zone bookkeeping.
    
    Fixes: 89d947561077 ("sd: Implement support for ZBC devices")
    Signed-off-by: ZHOU Jiaxiang <[email protected]>
    Reviewed-by: Damien Le Moal <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Martin K. Petersen (Oracle) <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

scsi: ufs: core: Keep internal commands dispatchable during error handling [+ + +]
Author: Stanley Jhu <[email protected]>
Date:   Sat Sep 12 21:16:25 2026 +0800

    scsi: ufs: core: Keep internal commands dispatchable during error handling
    
    commit b52d695d062095327b944acf7daabbc816ab319b upstream.
    
    Commit 08b12cda6c44 ("scsi: ufs: core: Switch to scsi_get_internal_cmd()")
    switched UFS internal commands to allocate requests on
    hba->host->pseudo_sdev->request_queue, which shares the host tagset with
    regular LUNs.
    
    During error recovery, ufshcd_err_handling_prepare() calls
    blk_mq_quiesce_tagset(&hba->host->tag_set), marking all queues in the
    tagset as quiesced, including pseudo_sdev->request_queue. When
    ufshcd_verify_dev_init() subsequently issues internal commands (e.g. NOP
    OUT UPIU) via blk_execute_rq(), blk_mq_run_hw_queue() skips running the
    quiesced queue, resulting in an unrecoverable circular wait deadlock.
    
    Keep quiescing the tagset and unquiesce the pseudo SCSI device on top of
    that, so internal commands stay dispatchable while the logical units remain
    quiesced. Re-quiesce the pseudo device before unquiescing the tagset so
    that quiesce_depth stays balanced.
    
    Clock scaling and ufshcd_pause_command_processing() are unaffected: they
    keep quiescing the whole tagset, internal commands included.
    
    Fixes: 08b12cda6c44 ("scsi: ufs: core: Switch to scsi_get_internal_cmd()")
    Cc: [email protected]
    Link: https://lore.kernel.org/all/[email protected]/
    Signed-off-by: Stanley Jhu <[email protected]>
    Reviewed-by: Bart Van Assche <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Martin K. Petersen (Oracle) <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

scsi: ufs: pltfrm: Add quirk for R-Car S4 lacking lanes-per-direction [+ + +]
Author: Geert Uytterhoeven <[email protected]>
Date:   Mon Sep 14 16:00:01 2026 +0200

    scsi: ufs: pltfrm: Add quirk for R-Car S4 lacking lanes-per-direction
    
    commit c9ee6511332687ea714ad8ab86a53cb837d86eea upstream.
    
    Since commit e72323f3b09f ("scsi: ufs: core: Configure only active lanes
    during link"), the following error is observed on R-Car S4:
    
        ufshcd-renesas e6860000.ufs: Tx lane mismatch [config,reported] [2,1]
        ufshcd-renesas e6860000.ufs: link startup failed -67
        ufshcd-renesas e6860000.ufs: error -ENOLINK: Initialization failed with error -67
        ufshcd-renesas e6860000.ufs: probe with driver ufshcd-renesas failed with error -67
    
    R-Car S4 has one UFS lane per direction, as described in section 152.1 of
    its hardware manual.  Without lanes-per-direction, the UFS platform driver
    defaults to two lanes.
    
    Previously, the core used PA_CONNECTEDRXDATALANES and
    PA_CONNECTEDTXDATALANES to configure the link without checking them against
    lanes-per-direction, so the missing property did not prevent
    initialization.
    
    While fixing the R-Car S4 DTS is the proper solution, doing only that would
    still break backwards compatibility with existing DTBs.  Hence add a quirk
    to let lanes-per-direction default to one on R-Car S4.
    
    Fixes: e72323f3b09f9c89 ("scsi: ufs: core: Configure only active lanes during link")
    Reported-by: Koichiro Den <[email protected]>
    Closes: https://lore.kernel.org/[email protected]
    Cc: [email protected] # 7.2+
    Signed-off-by: Geert Uytterhoeven <[email protected]>
    Link: https://patch.msgid.link/ae0cc2bd764e6dfffce99db3d8b44a55887c508c.1789394185.git.geert+renesas@glider.be
    Signed-off-by: Martin K. Petersen (Oracle) <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
sctp: avoid livelock while updating retransmit path [+ + +]
Author: Yiqi Sun <[email protected]>
Date:   Tue Sep 15 17:50:17 2026 +0800

    sctp: avoid livelock while updating retransmit path
    
    [ Upstream commit d2c31b837406395e576afeb25958c98e9938f3f6 ]
    
    sctp_assoc_update_retran_path() can loop forever when every remaining
    transport, including retran_path, is SCTP_UNCONFIRMED: the state check
    runs before the wraparound test, so the loop cannot observe that it has
    completed a full pass.
    
    Fix this by considering a transport only when it is not UNCONFIRMED,
    then checking whether the walk has returned to retran_path. This makes
    the full-pass termination independent of the transport state while
    preserving the existing fallback selection semantics.
    
    Also restore the NULL guard around the retran_path assignment. In the
    all-UNCONFIRMED case there is no eligible replacement transport, and
    installing NULL would leave later retransmit-path users and the debug
    print with a NULL path.
    
    Fixes: 4c47af4d5eb2 ("net: sctp: rework multihoming retransmission path selection to rfc4960")
    Signed-off-by: Yiqi Sun <[email protected]>
    Acked-by: Xin Long <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

sctp: discard the rest of the packet on a stale-cookie error [+ + +]
Author: Aohan Mei <[email protected]>
Date:   Mon Sep 21 17:37:04 2026 +0800

    sctp: discard the rest of the packet on a stale-cookie error
    
    commit 4498467a8af06cfa3d71cb04bd7c4170dec8f449 upstream.
    
    When an association is in COOKIE-ECHOED state and the peer sends a
    bundled [ERROR(Stale Cookie)][DATA] packet from one of its non-primary
    addresses, processing the ERROR chunk takes the non-fatal stale-cookie
    retry path sctp_sf_do_5_2_6_stale(), which queues
    SCTP_CMD_DEL_NON_PRIMARY while keeping the association alive.
    sctp_cmd_del_non_primary() removes every non-primary transport -
    including the very transport this packet arrived on, which is still
    referenced by the receive lookup and shared by all chunks of the
    packet via chunk->transport.
    
    sctp_assoc_rm_peer() does redirect asoc->peer.last_data_from away from
    the removed transport, but right afterwards the bundled DATA chunk
    makes sctp_assoc_bh_rcv() re-register
    asoc->peer.last_data_from = chunk->transport unconditionally, undoing
    the redirection with the just-removed transport.
    
    Once the packet is done, the receive reference is dropped and the
    transport is RCU-freed, while the surviving association keeps the
    dangling last_data_from.  A later FWD-TSN (or the delayed SACK timer)
    makes sctp_gen_sack() dereference it (->param_flags and friends), and
    sctp_make_sack()/sctp_outq_select_transport() may write to the freed
    object and link it into the live transport list.  This is a
    use-after-free triggerable by any malicious SCTP peer (or a local
    unprivileged user acting as one) with no capabilities required:
    
      BUG: KASAN: slab-use-after-free in sctp_do_sm+0x498a/0x5660
      Read of size 4 at addr ffff88800e1e356c by task poc/115
      Call Trace: sctp_do_sm <- sctp_assoc_bh_rcv <- sctp_inq_push <-
                  sctp_rcv <- ip_protocol_deliver_rcu <- ip_rcv
      Allocated: sctp_transport_new <- sctp_assoc_add_peer <-
                 sctp_process_init (INIT-ACK processing)
      Freed: kfree <- sctp_transport_destroy_rcu <- rcu_core
             (call_rcu queued by sctp_transport_put at end of sctp_rcv)
      The buggy address is located 364 bytes inside of freed 1024-byte
      region [ffff88800e1e3400, ffff88800e1e3800), cache kmalloc-1k
    
    Note that commit 03a9d10ecf71 ("sctp: drop a chunk if its transport
    was removed") only covers the window between the receive lookup and
    the chunk processing (e.g. an ASCONF DEL-IP racing the socket backlog);
    here the transport is removed *while* the packet is being processed,
    by an earlier chunk of the same packet, so the drop in sctp_inq_push()
    does not reach this path.  Verified with the bundled [ERROR(Stale
    Cookie)][DATA] + FWD-TSN reproducer: the KASAN report above still
    fires with that commit applied, and is gone with this patch on top.
    
    Fix it by discarding the rest of the packet on this path, as suggested
    by Xin.  After the stale-cookie ERROR has sent the association back to
    COOKIE-WAIT and removed the non-primary transports, the remaining
    chunks of the packet can only run against the restarted handshake
    while referencing the removed arrival transport through
    chunk->transport: besides the last_data_from registration above,
    sctp_cmd_setup_t2() and the sctp_make_*() reply builders would also
    copy that pointer into association-lifetime state that
    sctp_assoc_rm_peer() has already sanitized.  Let the peer retransmit
    them, in line with what sctp_inq_push() does for chunks whose
    transport was removed before processing.
    
    Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
    Suggested-by: Xin Long <[email protected]>
    Reported-by: TencentOS Corvus AI <[email protected]>
    Cc: [email protected]
    Signed-off-by: Aohan Mei <[email protected]>
    Acked-by: Xin Long <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

sctp: hold asoc or transport before mod_timer() in timer handlers [+ + +]
Author: Xin Long <[email protected]>
Date:   Mon Sep 21 14:03:45 2026 -0400

    sctp: hold asoc or transport before mod_timer() in timer handlers
    
    [ Upstream commit cae23ae3f7887a1cf8a75da38edcebeef695040f ]
    
    Take the association or transport reference before rearming a timer in the
    timer handlers.
    
    The existing code calls mod_timer() before taking the reference needed by
    the rearmed timer without holding the sock lock. This creates a race with
    timer cleanup: if the timer is deleted after mod_timer() returns but before
    the reference is taken, the cleanup path can drop the timer's reference and
    destroy the transport or association. The timer handler then takes a
    reference on the already freed object and eventually drops it, causing a
    refcount underflow.
    
    Hold the object before mod_timer() and drop the reference if mod_timer()
    reports that the timer was already pending in timer handlers. Apply the
    same ordering to the proto-unreachable path, which can rearm a transport
    timer outside the timer handlers without holding the sock lock.
    
    Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
    Reported-by: Tangxin Xie <[email protected]>
    Signed-off-by: Xin Long <[email protected]>
    Link: https://patch.msgid.link/c31b5e3ee2b7274e804f5eba2f21e2412e7eef7a.1790013825.git.lucien.xin@gmail.com
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
selftests/ftrace: Fix unique symbol check in kprobe_non_uniq_symbol.tc [+ + +]
Author: Sven Schnelle <[email protected]>
Date:   Wed Sep 9 11:29:53 2026 +0200

    selftests/ftrace: Fix unique symbol check in kprobe_non_uniq_symbol.tc
    
    [ Upstream commit d22c3e0088e85be8131f7a9283f759cdbb20726d ]
    
    The current regex also matches symbols in modules, which makes the
    test fail on s390 where name_show is present only once in the kernel,
    but also multiple times in modules:
    
    000001b1401cdc20 t name_show
    000001b0c05e6c40 t name_show    [mdev]
    000001b0c0495f30 t name_show    [i2c_core]
    
    Fix this by changing the regular expression to only match the function
    name.
    
    Link: https://lore.kernel.org/all/[email protected]/
    
    Fixes: 03b80ff8023a ("selftests/ftrace: Add new test case which checks non unique symbol")
    Signed-off-by: Sven Schnelle <[email protected]>
    Reviewed-by: Steven Rostedt <[email protected]>
    Signed-off-by: Masami Hiramatsu (Google) <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
selftests/nci: Fix out-of-bounds store on thread join [+ + +]
Author: Chris Gellermann <[email protected]>
Date:   Fri Sep 4 11:59:15 2026 +0200

    selftests/nci: Fix out-of-bounds store on thread join
    
    [ Upstream commit 6be581aeffc215bfc77939cd59902b0dbc4af23e ]
    
    The NCI test collects the exit status of its helper threads by passing
    the address of an int to pthread_join():
    
            int status;
            ...
            pthread_join(thread_t, (void **) &status);
    
    pthread_join() stores a void pointer to the memory location. On 64-bit
    systems, a void pointer is wider than an int, so the store overruns the
    4 bytes of space allocated on the stack for the integer and corrupts the
    adjacent stack. On our CHERI system, this caused a fault due to a
    capability bounds violation.
    
    Fix this by introducing a helper that joins a thread through a void
    pointer and converts the result back to an integer, which is what the
    helper threads return.
    
    While here, also fix the logic in disconnect_tag() if the helper thread
    creation failed. Previously, it would have joined a thread that was
    never created when pthread_create() failed.
    
    Fixes: f595cf1242f3 ("selftests: Add nci suite")
    Signed-off-by: Chris Gellermann <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
selftests: cgroup: give the O_TMPFILE open in get_temp_fd() a mode [+ + +]
Author: Eva Kurchatova <[email protected]>
Date:   Wed Sep 16 23:44:31 2026 +0300

    selftests: cgroup: give the O_TMPFILE open in get_temp_fd() a mode
    
    [ Upstream commit c774ec8f0a5d02a06d34c27f5a7de7e333b91265 ]
    
    O_TMPFILE, like O_CREAT, needs the third argument. Without it glibc
    refuses the call at compile time as soon as fortification is on:
    
      In function 'open',
          inlined from 'get_temp_fd' at test_memcontrol.c:33:9:
      /usr/include/bits/fcntl2.h:52:11: error: call to '__open_missing_mode'
        declared with attribute error: open with O_CREAT or O_TMPFILE in
        second argument needs 3 arguments
    
    The fortify checks take effect only once the compiler optimises, and
    cgroup/Makefile builds with "-Wall -pthread" alone, so this goes
    unnoticed in a plain build. Building the tests with the flags
    distributions commonly use, -O2 -D_FORTIFY_SOURCE=3, loses
    test_memcontrol entirely.
    
    Fixes: 84092dbcf901 ("selftests: cgroup: add memory controller self-tests")
    Signed-off-by: Eva Kurchatova <[email protected]>
    Signed-off-by: Tejun Heo <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

selftests: nci: Correct pthread_create return value check [+ + +]
Author: Lei Zhu <[email protected]>
Date:   Wed Jul 29 15:24:26 2026 +0800

    selftests: nci: Correct pthread_create return value check
    
    [ Upstream commit 3d8afc5243ea2ee803d98e69eb4a01748167ac1b ]
    
    The pthread_create() functions returns 0 on success and a positive value on
    failure. Modify the return value check to correctly detect failure cases.
    
    Fixes: 72696bd8a09d ("selftests: nci: Extract the start/stop discovery function")
    Signed-off-by: Lei Zhu <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

selftests: nci: Fix uninitialized family ID on missing attribute [+ + +]
Author: Chaithanya Lagisetty <[email protected]>
Date:   Tue Sep 1 07:06:18 2026 +0000

    selftests: nci: Fix uninitialized family ID on missing attribute
    
    [ Upstream commit eda518d2cdb6074a0bcdfa06af291616bcb5c421 ]
    
    get_family_id() walks the generic netlink CTRL_CMD_GETFAMILY reply
    looking for the CTRL_ATTR_FAMILY_ID attribute and returns the parsed
    value in the local variable "id". If the reply does not carry that
    attribute, the parsing loop never assigns "id" and the function returns
    an indeterminate stack value, which the caller stores in self->fid and
    uses for subsequent netlink requests.
    
    Initialize "id" to 0 so a missing attribute yields a deterministic
    (invalid) family ID instead of a garbage value.
    
    Fixes: f595cf1242f3 ("selftests: Add nci suite")
    Signed-off-by: Chaithanya Lagisetty <[email protected]>
    Reviewed-by: Hangbin Liu <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: David Heidelberg <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
smb: client: clean up failed cached directory opens [+ + +]
Author: Zihan Xi <[email protected]>
Date:   Wed Sep 16 15:29:27 2026 +0000

    smb: client: clean up failed cached directory opens
    
    commit d2ff5fb93ea83034025850266b5eed391f96b825 upstream.
    
    open_cached_dir() sends CREATE and QUERY_INFO as a compound request. If
    the CREATE succeeds but a later command returns an error, the function
    must retain the CREATE FID so common cleanup can issue SMB2_close(). It
    also must not treat a response error as a valid CREATE.
    
    Validate the CREATE response before using its fields, record the FIDs, and
    mark the handle open before handling errors from later compound commands.
    Move the -EREMCHG reconnect handling before response validation so a
    missing response does not hide the reconnect request. Count the handle
    when it is marked open; confirmed close responses decrement the counter,
    while existing close retry behavior remains best effort on transport
    failures.
    
    Fixes: b0f6df737a1c ("cifs: cache FILE_ALL_INFO for the shared root handle")
    Cc: [email protected]
    Reported-by: Vega <[email protected]>
    Assisted-by: LLM
    Co-developed-by: Luxing Yin <[email protected]>
    Signed-off-by: Luxing Yin <[email protected]>
    Signed-off-by: Zihan Xi <[email protected]>
    Tested-by: Frank Sorenson <[email protected]>
    Signed-off-by: Paulo Alcantara <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

smb: client: close completed creates on compound wait errors [+ + +]
Author: Zihan Xi <[email protected]>
Date:   Wed Sep 16 15:29:28 2026 +0000

    smb: client: close completed creates on compound wait errors
    
    commit 6c5c547f037bc18f0b8d0b5db5a648f8f630ce85 upstream.
    
    compound_send_recv() waits for responses in order. If a later wait is
    interrupted, or if a later MID fails during response synchronization, an
    earlier CREATE may already have opened a remote handle. The earlier mid
    is then released without invoking handle_cancelled_mid(), leaving the
    remote handle open because no FID was copied to the caller.
    
    Mark completed earlier mids as cancelled when a compound wait or MID
    synchronization aborts. Keep their response buffers attached while the
    MIDs are synchronized, and transfer them only after synchronization of
    the processed responses, so the release path can inspect successful
    CREATE responses and queue SMB2_close() after a later failure. Account for
    a remote open only after the close work is allocated and before it is
    queued, since the caller has not yet updated num_remote_opens. Mark the
    create+close compound used by smb2_unlink() so it is not closed again.
    Non-CREATE responses and compounds that already include a close keep their
    existing behavior.
    
    Fixes: e0bba0b85481 ("cifs: add compound_send_recv()")
    Cc: [email protected]
    Reported-by: Vega <[email protected]>
    Assisted-by: LLM
    Co-developed-by: Luxing Yin <[email protected]>
    Signed-off-by: Luxing Yin <[email protected]>
    Signed-off-by: Zihan Xi <[email protected]>
    Tested-by: Frank Sorenson <[email protected]>
    Signed-off-by: Paulo Alcantara <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

smb: client: close handle after create-context parsing failure [+ + +]
Author: Zihan Xi <[email protected]>
Date:   Wed Sep 16 15:29:26 2026 +0000

    smb: client: close handle after create-context parsing failure
    
    commit 566820af017e81497fb5e9d3ad6e7ffe2828bc8b upstream.
    
    SMB2_open() accounts a successful CREATE response as a remote open before
    parsing its create contexts.  If smb2_parse_contexts() rejects malformed
    context data, SMB2_open() returns without closing the handle, leaving the
    server-side handle open and num_remote_opens elevated.
    
    Close the handle after a post-CREATE context parsing failure so the error
    path releases the remote resource and balances the open count.
    
    Fixes: af1689a9b770 ("smb: client: fix potential OOBs in smb2_parse_contexts()")
    Cc: [email protected]
    Reported-by: Vega <[email protected]>
    Assisted-by: LLM
    Co-developed-by: Luxing Yin <[email protected]>
    Signed-off-by: Luxing Yin <[email protected]>
    Signed-off-by: Zihan Xi <[email protected]>
    Tested-by: Frank Sorenson <[email protected]>
    Signed-off-by: Paulo Alcantara <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

smb: client: delete compound mids on send failure before unlock [+ + +]
Author: Adarsh Das <[email protected]>
Date:   Sat Sep 19 16:27:32 2026 +0530

    smb: client: delete compound mids on send failure before unlock
    
    commit 8f6f8a48399f82f4a83f7b9f25b9707e0062a4f9 upstream.
    
    When sending a compound request fails, smb_send_rqst() kicks off a
    reconnect. compound_send_recv() still has those mids on pending_mid_q,
    but it unlocks the server without removing them first.
    
    During reconnect, cifs_abort_connection() walks pending_mid_q and runs
    each mid callback. With no response yet, those callbacks return credits
    and drop in_flight. Then compound_send_recv()'s send-error path returns
    the same credits again. in_flight ends up decremented twice and
    smb2_add_credits() WARNs.
    
    syzbot hits this during SMB2_negotiate when the socket send fails.
    
    cifs_call_async() already calls delete_mid() before unlock on send
    failure. Do the same for compound chains and set cancelled_mid[] so the
    out: path does not delete them again.
    
    Fixes: ee258d79159a ("CIFS: Move credit processing to mid callbacks for SMB3")
    Reported-by: [email protected]
    Closes: https://lore.kernel.org/r/[email protected]
    Tested-by: [email protected]
    Assisted-by: LLM
    Cc: [email protected]
    Signed-off-by: Adarsh Das <[email protected]>
    Signed-off-by: Paulo Alcantara <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

smb: client: fix create context out-of-bounds reads [+ + +]
Author: Zihan Xi <[email protected]>
Date:   Wed Sep 16 15:29:24 2026 +0000

    smb: client: fix create context out-of-bounds reads
    
    commit 67f4c1c6a1b51e203d986779299824d1c2c590a6 upstream.
    
    smb2_parse_contexts() validates the complete create-context area but
    does not limit each record to its Next field before dispatching it.  A
    malformed chain can therefore expose bytes beyond the current context to
    a handler.  The QFid handler also used a full response-structure cast
    although it only reads DiskFileId.
    
    The SMB2/SMB3 lease parsers made the same layout assumption: they read
    LeaseState and LeaseFlags at canonical offsets rather than at
    DataOffset.  A valid non-canonical DataOffset could therefore yield
    unrelated in-bounds data, while a short DataLength was still accepted.
    
    Limit each context to its Next value, reject offsets before the context
    header, and reject malformed chains.  Bound the name range by the current
    context and do not dispatch a known handler when DataLength is zero.  Read
    the QFid DiskFileId only when the context data covers that field.  Parse the
    lease context from DataOffset and require DataLength to match the v1 or v2
    lease_context size used by ksmbd.  A size mismatch skips lease parsing
    without failing the open.
    
    Fixes: b8c32dbb0deb ("CIFS: Request SMB2.1 leases")
    Fixes: f047390a097e ("CIFS: Add create lease v2 context for SMB3")
    Fixes: 89a5bfa350fa ("smb3: optimize open to not send query file internal info")
    Cc: [email protected]
    Reported-by: Vega <[email protected]>
    Assisted-by: LLM
    Co-developed-by: Luxing Yin <[email protected]>
    Signed-off-by: Luxing Yin <[email protected]>
    Signed-off-by: Zihan Xi <[email protected]>
    Tested-by: Frank Sorenson <[email protected]>
    Signed-off-by: Paulo Alcantara <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

smb: client: preserve create-context parsing errors [+ + +]
Author: Zihan Xi <[email protected]>
Date:   Wed Sep 16 15:29:29 2026 +0000

    smb: client: preserve create-context parsing errors
    
    commit 2e828035d5d736d904c238ae7ec3f77c0d6270bf upstream.
    
    smb2_compound_op() saves the result from compound_send_recv() in
    tmp_rc. For SMB2_OP_OPEN_QUERY it then parses the CREATE contexts, but
    the final assignment of rc from tmp_rc discards a parsing error. A
    malformed create-context response can therefore be reported as
    successful to smb2_query_path_info().
    
    Keep a create-context parsing error in tmp_rc so it survives per-command
    response processing and is returned to the caller.
    
    Fixes: b07687edee99 ("cifs: Improve SMB2+ stat() to work also without FILE_READ_ATTRIBUTES")
    Cc: [email protected]
    Reported-by: Vega <[email protected]>
    Assisted-by: LLM
    Co-developed-by: Luxing Yin <[email protected]>
    Signed-off-by: Luxing Yin <[email protected]>
    Signed-off-by: Zihan Xi <[email protected]>
    Tested-by: Frank Sorenson <[email protected]>
    Signed-off-by: Paulo Alcantara <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

smb: client: use finish_no_open() for non-regular inodes [+ + +]
Author: Namjae Jeon <[email protected]>
Date:   Wed Sep 23 20:25:49 2026 +0900

    smb: client: use finish_no_open() for non-regular inodes
    
    commit e66cf1625ec4a3fe68346119f371def713fd0a4d upstream.
    
    An O_CREAT open can find an existing symlink or another non-regular
    inode. cifs_atomic_open() calls finish_open() on it and attaches a
    cifsFileInfo. Symlink inodes have no CIFS release operation, so the
    dentry reference held by cifsFileInfo is leaked. FMODE_OPENED also
    prevents the VFS from following the symlink.
    
    Track whether cifs_do_create() returned an open server handle. For
    non-regular inodes, close the handle if present, remove the pending
    open, and call finish_no_open() so the VFS can continue the lookup.
    Do not set FMODE_CREATED unless a regular file was opened. For
    O_NOFOLLOW with __O_REGULAR, return -ELOOP before the VFS's
    -EFTYPE check.
    
    Defer closing a legacy POSIX handle on a non-regular inode until
    after inode lookup. This avoids closing it again if lookup fails.
    
    Fixes: d2c127197dfc ("cifs: implement i_op->atomic_open()")
    Signed-off-by: Namjae Jeon <[email protected]>
    Signed-off-by: Paulo Alcantara <[email protected]>
    Cc: [email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

smb: client: validate POSIX create context length [+ + +]
Author: Zihan Xi <[email protected]>
Date:   Wed Sep 16 15:29:25 2026 +0000

    smb: client: validate POSIX create context length
    
    commit fa2e9900dd2a3f5a1e7ef5a8c5e8d435feedbfcc upstream.
    
    parse_posix_ctxt() reads the fixed nlink, reparse_tag, and mode fields
    before checking that the POSIX create context contains them.  A short
    context can pass the generic checks and still make these fixed-width
    reads run past its declared data.
    
    The current in-tree smb2_open_file() path passes a NULL posix pointer,
    so this handler is not reached on the ordinary open path.  Still require
    the POSIX data to cover all three fields before reading them because the
    helper performs those unguarded reads.  Keep the existing soft-failure
    behavior so malformed optional metadata does not fail the open.
    
    Fixes: 69dda3059e7a ("cifs: add SMB2_open() arg to return POSIX data")
    Cc: [email protected]
    Reported-by: Vega <[email protected]>
    Assisted-by: LLM
    Co-developed-by: Luxing Yin <[email protected]>
    Signed-off-by: Luxing Yin <[email protected]>
    Signed-off-by: Zihan Xi <[email protected]>
    Tested-by: Frank Sorenson <[email protected]>
    Signed-off-by: Paulo Alcantara <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
squashfs: Add dictionary size range check to prevent shift-out-of-bounds [+ + +]
Author: Ran Hongyun <[email protected]>
Date:   Mon Jul 13 19:55:25 2026 +0800

    squashfs: Add dictionary size range check to prevent shift-out-of-bounds
    
    [ Upstream commit 1f7745fb3580152ca902ef181b605f33cabfb1d0 ]
    
    When an abnormal SquashFS image (COMP_OPTS flag is 1 but dictionary size
    is 0) is mounted, and performs shift operations using dictionarysize, the
    shift exponent is -1, causing a shift-out-of-bounds.
    
    Detail as below:
    squashfs_comp_opts(msblk, buffer, length)
      squashfs_xz_comp_opts()
        if (comp_opts)
          n = ffs(opts->dict_size) - 1;<----opts->dict_size=0, n=-1
          if (opts->dict_size != (1 << n) && opts->dict_size !=
                    (1 << n) + (1 << (n + 1))) <----shift-out-of-bounds
    
    Fix it by adding a dictionary size range check before the shift operation.
    
    Fixes: ff750311d30a ("Squashfs: add compression options support to xz decompressor")
    Signed-off-by: Ran Hongyun <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Reviewed-by: Phillip Lougher <[email protected]>
    Reviewed-by: Zhihao Cheng <[email protected]>
    Signed-off-by: Christian Brauner (Amutable) <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
super: convert s_count to refcount_t s_passive [+ + +]
Author: Christian Brauner <[email protected]>
Date:   Tue Sep 29 09:52:28 2026 -0400

    super: convert s_count to refcount_t s_passive
    
    [ Upstream commit 3ec9800c2d33c783dd3b27d4cc3bb22b9385f828 ]
    
    The superblock carries two counters: s_active, the active reference
    count that keeps the filesystem usable, and s_count, the passive
    reference count that merely keeps the structure itself alive. Turn the
    passive count into a refcount_t and rename it to s_passive to make the
    pairing with s_active obvious.
    
    Everything is still serialized by sb_lock, so there is no functional
    change; the conversion buys the usual refcount_t saturation and
    underflow checking. The following patches start dropping passive
    references without holding sb_lock and make the device-to-superblock
    table hold one passive reference per registered entry, which a plain
    integer cannot support.
    
    Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-2-7df6b864028e@kernel.org
    Reviewed-by: Jan Kara <[email protected]>
    Signed-off-by: Christian Brauner (Amutable) <[email protected]>
    Stable-dep-of: 2d2a2d7aa987 ("super: make iterate_supers_type() deletion-safe")
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

super: make iterate_supers_type() deletion-safe [+ + +]
Author: Christian Brauner <[email protected]>
Date:   Tue Sep 29 09:52:30 2026 -0400

    super: make iterate_supers_type() deletion-safe
    
    [ Upstream commit 2d2a2d7aa98741b58f54cacc99b52024e4d865f9 ]
    
    iterate_supers_type() drops sb_lock while invoking the callback and keeps
    only a passive reference to the current superblock. That reference keeps
    the object allocated, but does not keep its s_instances node linked.
    
    After the iterator releases s_umount, final teardown can unlink the current
    s_instances node. The iterator then advances through a reinitialized node.
    With the current hlist it stops without visiting the remaining superblocks.
    The unlink moved from generic_shutdown_super() to kill_super_notify(), but
    the cursor lifetime has been unsafe since the helper was introduced.
    
    The CIFS DFS lookup can consequently miss a matching superblock and return
    -EINVAL.
    
    Move removal from fs_supers to put_super(), alongside removal from
    super_blocks, so a passive reference keeps both list nodes linked. Keep
    the filesystem module reference until then, since unlinking s_instances
    may touch type->fs_supers.
    
    Make sget_fc() skip SB_DEAD superblocks before invoking test(), and set
    SB_DEAD under sb_lock to serialize with those callbacks. This allows
    kernfs to free its private information after kill_anon_super() returns.
    Keep matching SB_DYING superblocks until SB_DEAD is set so concurrent
    mounts still wait for teardown before retrying.
    
    Fixes: 43e15cdbefea ("new helper: iterate_supers_type()")
    Reported-by: Karl Mehltretter <[email protected]>
    Closes: https://lore.kernel.org/r/[email protected]
    Suggested-by: Jan Kara <[email protected]>
    Cc: [email protected]
    Tested-by: Karl Mehltretter <[email protected]>
    [kmehltretter: supplied the commit message]
    Signed-off-by: Karl Mehltretter <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Reviewed-by: Jan Kara <[email protected]>
    Signed-off-by: Christian Brauner (Amutable) <[email protected]>
    [ adjusted kill_super_notify() context for missing device-to-superblock table cleanup. ]
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

super: take lock after last reference count [+ + +]
Author: Christian Brauner <[email protected]>
Date:   Tue Sep 29 09:52:29 2026 -0400

    super: take lock after last reference count
    
    [ Upstream commit 9c486f28994fdfc1a83f5f60129402cf4379957b ]
    
    __put_super() required the caller to hold sb_lock, so put_super()
    wrapped it. The per-device superblock table introduced later drops its
    passive references from contexts that do not hold sb_lock, so make
    put_super() self-locking: drop the count first and take sb_lock only for
    the final list_del.
    
    With the count now dropped outside sb_lock a superblock can briefly sit
    on @super_blocks with s_passive == 0 before it is unlinked, so the list
    walkers (__iterate_supers(), iterate_supers_type(), user_get_super())
    switch to refcount_inc_not_zero() and skip it.
    
    Link: https://patch.msgid.link/20260616-work-super-bdev_holder_global-v2-3-7df6b864028e@kernel.org
    Reviewed-by: Jan Kara <[email protected]>
    Signed-off-by: Christian Brauner (Amutable) <[email protected]>
    Stable-dep-of: 2d2a2d7aa987 ("super: make iterate_supers_type() deletion-safe")
    Signed-off-by: Sasha Levin <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
tcp: fix use-after-free of retransmit_skb_hint in tcp_send_synack() [+ + +]
Author: Yilin Zhang <[email protected]>
Date:   Thu Sep 24 12:49:00 2026 +0800

    tcp: fix use-after-free of retransmit_skb_hint in tcp_send_synack()
    
    [ Upstream commit fe99bbeee5c5dbd3abc30721a8079ced59649d97 ]
    
    When tcp_send_synack() replaces the cloned SYN skb at the head of the
    retransmit queue with a copy, it frees the original with
    tcp_rtx_queue_unlink_and_free() and only repairs tp->highest_sack.
    tp->retransmit_skb_hint keeps pointing at the freed
    skbuff_fclone_cache object.
    
    The dangling hint is read in tcp_verify_retransmit_hint() and used as
    the root of the rbtree walk in tcp_xmit_retransmit_queue().  An
    unprivileged TFO client (sendmsg(MSG_FASTOPEN)) can arm the hint with
    an attacker-supplied ICMP fragmentation-needed message, after which a
    simultaneous open frees the armed SYN skb:
    
      BUG: KASAN: slab-use-after-free in tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316)
      Read of size 4 at addr ffff88800604d928 by task swapper/1/0
      Call Trace:
       tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316)
       tcp_simple_retransmit (net/ipv4/tcp_input.c:3158)
       tcp_v4_err (net/ipv4/tcp_ipv4.c:587)
    
    Sync the hint to the copy.
    
    Fixes: c31b70c9968f ("tcp: Add logic to check for SYN w/ data in tcp_simple_retransmit")
    Reported-by: Kimi Security Team <[email protected]>
    Tested-by: Weiming Shi <[email protected]>
    Signed-off-by: Yilin Zhang <[email protected]>
    Reviewed-by: Eric Dumazet <[email protected]>
    Link: https://patch.msgid.link/8a9dff4063a2745653b7e88ceb745d75efa16e68.1790224474.git.yilinzhang@moonshot.ai
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

tcp: prevent collapsing skbs across boundary in rtx queue [+ + +]
Author: Willem de Bruijn <[email protected]>
Date:   Thu Sep 24 11:44:12 2026 -0400

    tcp: prevent collapsing skbs across boundary in rtx queue
    
    commit fc6d80eb504458d6416b75a94188b268c95c6533 upstream.
    
    tcp_write_collapse_fence() sets TCP_SKB_CB(skb)->eor = 1 on
    tcp_write_queue_tail(sk) to prevent skbs queued after a switch to
    device encryption from being collapsed into earlier skbs.
    
    The fence is a no-op if all earlier data has already been transmitted
    when the switch happens: sk->sk_write_queue is empty. The not yet
    acknowledged earlier skbs wait in sk->tcp_rtx_queue with eor 0.
    
    On a subsequent retransmit or SACK shift, tcp_retrans_try_collapse() or
    tcp_shift_skb_data() can then merge an skb queued after the switch into
    one queued before it.
    
    Both users of the fence are affected:
    
    - psp: devices only encrypt skbs with skb->decrypted set. The merged skb
      keeps decrypted = 0 from the earlier skb, so merged data sent after
      psp_sock_assoc_set_tx() is retransmitted in cleartext.
    
    - tls device offload: the merged skb straddles the start marker set in
      tls_set_device_offload(). The software fallback (fill_sg_in() returns
      -EINVAL) and the mlx5, nfp and funeth drivers cannot handle such an
      skb and drop it. Every retransmit rebuilds the same skb, so the
      connection stalls.
    
    Fix this in two places, for defense in depth:
    
    1. Fall back to tcp_rtx_queue_tail(sk) in tcp_write_collapse_fence()
       when tcp_write_queue_tail(sk) is NULL.
    
    2. Check !skb_cmp_decrypted(to, from) in tcp_skb_can_collapse(), as
       tcp_skb_can_collapse_rx() does on receive. skb_shift(), which both
       collapse paths call, already has a DEBUG_NET_WARN_ON_ONCE() for this
       condition.
    
    Fixes: e8f69799810c ("net/tls: Add generic NIC offload infrastructure")
    Cc: [email protected]
    Signed-off-by: Willem de Bruijn <[email protected]>
    Reviewed-by: Eric Dumazet <[email protected]>
    Reviewed-by: Daniel Zahka <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

tcp: Skip cond_resched() in inet_csk_listen_stop() under BPF context [+ + +]
Author: Jiayuan Chen <[email protected]>
Date:   Thu Sep 10 19:27:28 2026 +0800

    tcp: Skip cond_resched() in inet_csk_listen_stop() under BPF context
    
    [ Upstream commit eaab8cab451b9502ce224cd202550375b894a467 ]
    
    bpf_sock_destroy() runs from the tcp iterator, under rcu_read_lock(). If
    the sock is a listener that still has children in its accept queue,
    tcp_abort() ends up in inet_csk_listen_stop() and the cond_resched()
    there trips the debug check:
    
    BUG: sleeping function called from invalid context at net/ipv4/inet_connection_sock.c:1523
    in_atomic(): 0, irqs_disabled(): 0, non_block: 0, pid: 628, name: test_progs
    preempt_count: 0, expected: 0
    RCU nest depth: 1, expected: 0
    locks held by test_progs/628: 3, last CPU#3:
     #0: ffff8881158cee18 (&p->lock){+.+.}-{4:4}, at: bpf_seq_read+0x56/0x1210
     #1: ffff8881106bb858 (sk_lock-AF_INET6){+.+.}-{0:0}, at: bpf_iter_tcp_seq_show+0x32b/0x4b0
     #2: ffffffffb435af20 (rcu_read_lock){....}-{1:3}, at: bpf_iter_run_prog+0x46b/0xde0
    CPU: 3 UID: 0 PID: 628 Comm: test_progs Tainted: G        W           7.2.0+ #65 PREEMPT
    Tainted: [W]=WARN
    Call Trace:
     <TASK>
     dump_stack_lvl+0xc1/0xf0
     dump_stack+0x10/0x20
     __might_resched+0x3d2/0x610
     inet_csk_listen_stop+0x7b/0xbf0
     tcp_abort+0x23b/0x3b0
     bpf_sock_destroy+0xfc/0x140
     bpf_prog_448133d24601754f_iter_tcp6_server+0x81/0x8a
     bpf_iter_run_prog+0x538/0xde0
     bpf_iter_tcp_seq_show+0x26b/0x4b0
     bpf_seq_read+0x424/0x1210
     vfs_read+0x197/0xe40
     ksys_read+0x119/0x240
     __x64_sys_read+0x72/0xc0
     x64_sys_call+0x647/0x27e0
     do_syscall_64+0xe5/0x610
     entry_SYSCALL_64_after_hwframe+0x76/0x7e
    RIP: 0033:0x7fad39b28aca
    RSP: 002b:00007ffc381c61c0 EFLAGS: 00000246 ORIG_RAX: 0000000000000000
    RAX: ffffffffffffffda RBX: 00007ffc381c6a88 RCX: 00007fad39b28aca
    RDX: 0000000000000032 RSI: 00007ffc381c6250 RDI: 0000000000000014
    RBP: 00007ffc381c61e0 R08: 0000000000000000 R09: 0000000000000000
    R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000003
    R13: 0000000000000000 R14: 000055f077c1bbb0 R15: 00007fad3a0f3000
     </TASK>
    
    The commit that added the kfunc already guards lock_sock() in tcp_abort()
    and udp_abort() with has_current_bpf_ctx(), but missed the listener path.
    Do the same for the cond_resched(). The loop runs inside the iterator's
    rcu_read_lock(), it must not reschedule or report a quiescent state there.
    
    Fixes: 4ddbcb886268 ("bpf: Add bpf_sock_destroy kfunc")
    Signed-off-by: Jiayuan Chen <[email protected]>
    Link: https://lore.kernel.org/r/[email protected]
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
tg3: clean up PHYLIB resources on probe failure [+ + +]
Author: Myeonghun Pak <[email protected]>
Date:   Thu Sep 17 14:33:36 2026 -0400

    tg3: clean up PHYLIB resources on probe failure
    
    [ Upstream commit a92e1a412c53dc0d9ad639e7abf8b3fc70a5b6ad ]
    
    tg3_get_invariants() can register an MDIO bus and connect a PHY for
    USE_PHYLIB devices. If tg3_init_one() later fails, its common error path
    releases the mappings and netdev without undoing those PHYLIB resources.
    
    Disconnect the PHY and unregister the MDIO bus before the remaining
    teardown. Guard PHY cleanup with USE_PHYLIB to match tg3_phy_init(), and
    call tg3_mdio_fini() unconditionally to match tg3_mdio_init(). The existing
    IS_CONNECTED and MDIOBUS_INITED flags make both helpers safe when
    initialization only completed partially.
    
    This issue was identified during our ongoing static-analysis research while
    reviewing kernel code.
    
    Fixes: 158d7abdae85 ("tg3: Add mdio bus registration")
    Assisted-by: OpenAI:GPT-5.6
    Co-developed-by: Ijae Kim <[email protected]>
    Signed-off-by: Ijae Kim <[email protected]>
    Signed-off-by: Myeonghun Pak <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

tg3: use random MAC address when tg3_get_device_address fails [+ + +]
Author: Ivan Delalande <[email protected]>
Date:   Fri Sep 18 15:47:15 2026 -0700

    tg3: use random MAC address when tg3_get_device_address fails
    
    [ Upstream commit 4eb3f195ef08c5acaed87958297e41cc49588dde ]
    
    Some of the tg3 NICs we use (BCM57762) reset the SRAM MAC address to the
    placeholder address on link flaps, tg3_chip_reset, etc. We've typically
    fixed it from userspace, but since e4c00ba7274b ("tg3: replace
    placeholder MAC address with device property") was merged, tg3 just
    fails probe as we don't have a way to get it through the generic
    device_get_mac_address infrastructure as fallback on our systems.
    
    Make the driver assign a random address in this condition instead of
    being fatal for probe. Set deferred_probe_reason through dev_warn_probe
    if the address isn't yet available from the provider.
    
    Fixes: e4c00ba7274b ("tg3: replace placeholder MAC address with device property")
    Suggested-by: Jakub Kicinski <[email protected]>
    Link: https://lore.kernel.org/netdev/[email protected]/
    Signed-off-by: Ivan Delalande <[email protected]>
    Link: https://patch.msgid.link/20260918224715.GA654128@visor
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
thermal: gov_step_wise: Fix stale mitigation vote with non-zero lower bounds [+ + +]
Author: Manaf Meethalavalappu Pallikunhi <[email protected]>
Date:   Tue Sep 22 17:49:02 2026 +0530

    thermal: gov_step_wise: Fix stale mitigation vote with non-zero lower bounds
    
    [ Upstream commit ec0d89150a9381d591344a9f6f5428655c227a7f ]
    
    When two or more thermal zones bind to a common cooling device and one zone
    uses a non-zero instance->lower value, there is a bug where the instance
    holds a stale mitigation vote even after its trip is cleared.
    
    Problem scenario:
    - thermal-zone1: Trip at 50°C, cooling-map with lower=0
    - thermal-zone2: Trip at 55°C, cooling-map with lower=2
    - Both zones share the same cooling device (e.g., CPU)
    
    Issue flow:
    1. Both trips trigger, zone1 requests state 5, zone2 also mitigates
    2. Zone2 trip clears (temp < 53°C due to hysteresis)
    3. When throttle=false and trend=THERMAL_TREND_DROPPING:
       - Current code checks: if (cur_state <= instance->lower)
         return THERMAL_NO_TARGET
       - Since cur_state (5) > instance->lower (2),
         it returns instance->lower (2)
       - This is the BUG where it returns instance->lower even though
         trip is cleared
    4. Zone2's passive polling stops (tz->passive reaches 0) - no more updates
       for zone2
    5. Zone2's stale vote of 2 persists indefinitely
    6. Even when zone1 wants to reduce cooling to state, the cooling device
       cannot go below state 2 due to zone2's stale vote
    
    When a trip is cleared (throttle == false), always return THERMAL_NO_TARGET
    instead of instance->lower. Remove the unnecessary check comparing
    cur_state with instance->lower. Since passive polling is already
    deactivated when the trip is cleared, the instance should always be
    deactivated regardless of its current cooling state. This ensures that
    instances with non-zero lower bounds do not retain stale mitigation votes
    after their trips are cleared.
    
    Fixes: 042a3d80f118 ("thermal: core: Move passive polling management to the core")
    Signed-off-by: Manaf Meethalavalappu Pallikunhi <[email protected]>
    Link: https://patch.msgid.link/20260922-step_wise_multi_zone_stale_vote_fix-v1-1-789f68dab229@oss.qualcomm.com
    Signed-off-by: Rafael J. Wysocki <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
tipc: Fix a data race on mon->peer_cnt in mon_timeout() [+ + +]
Author: Ginger Li <[email protected]>
Date:   Tue Sep 22 16:09:09 2026 +0800

    tipc: Fix a data race on mon->peer_cnt in mon_timeout()
    
    [ Upstream commit 8e1937fed6738460554ec123c64839e2445e7d53 ]
    
    mon_timeout() evaluates dom_size(mon->peer_cnt) before it takes mon->lock,
    while mon->peer_cnt is updated under that lock by tipc_mon_add_peer() and
    tipc_mon_remove_peer().  The value can therefore be stale, and the decision
    whether the local domain has to be recomputed can be based on an outdated
    member count.
    
    Read mon->peer_cnt inside the write_lock_bh(&mon->lock) protected region.
    
    Fixes: 35c55c9877f8 ("tipc: add neighbor monitoring framework")
    Signed-off-by: Ginger Li <[email protected]>
    Reviewed-by: Tung Nguyen <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

tipc: reject invalid and unexpected GRP_ACK_MSG to prevent bc_ackers underflow [+ + +]
Author: Eric Dumazet <[email protected]>
Date:   Sun Sep 13 04:42:33 2026 +0000

    tipc: reject invalid and unexpected GRP_ACK_MSG to prevent bc_ackers underflow
    
    commit 99cc2a62e07a44a22254d7beca9ef1f8ad886d0d upstream.
    
    Commit 48a5fe38772b ("tipc: fix bc_ackers underflow on duplicate
    GRP_ACK_MSG") rejected duplicate/stale ACKs in tipc_group_proto_rcv()
    by returning early when less_eq(acked, m->bc_acked).
    
    However, that check remains incomplete in two ways:
    
    1. When grp->bc_ackers is zero (e.g. on a quiet group, when replicast
       ACKs were not requested, or after all expected members have already
       acknowledged), an unexpected GRP_ACK_MSG with acked > m->bc_acked
       passes less_eq() and unconditionally decrements grp->bc_ackers.
       Because bc_ackers is a u16, this wraps to 65535, causing
       tipc_group_bc_cong() to permanently report congestion and blocking
       all future group broadcasts on the socket.
    
    2. During an active broadcast round (grp->bc_ackers > 0), the sender
       transmits packet S and advances grp->bc_snd_nxt to S + 1. Receivers
       increment their expected counter to S + 1 upon consuming packet S,
       so the only valid ACK value for the current round is strictly
       acked == grp->bc_snd_nxt.
    
       However, tipc_group_update_bc_members() initializes each member's
       m->bc_acked to prev = grp->bc_snd_nxt - 1 (S - 1 before increment).
       This leaves a 2-sequence gap (S - 1 to S + 1) in sequence space.
       An incoming ACK is therefore neither rejected as duplicate nor
       prevented from decrementing grp->bc_ackers if an unexpected or stale
       value (such as S) is received. A member sending acked = S followed
       by acked = S + 1 could decrement grp->bc_ackers twice in the same
       round, prematurely clearing bc_ackers or underflowing it.
    
    Fix this by:
    - Dropping GRP_ACK_MSG immediately if grp->bc_ackers is zero.
    - Requiring acked == grp->bc_snd_nxt and rejecting duplicates where
      m->bc_acked == acked. Because replicast broadcast rounds are strictly
      sequential, only grp->bc_snd_nxt can be acknowledged, and each member
      can acknowledge at most once per round.
    
    Note that a related pre-existing issue in tipc_group_delete_member()
    (where grp->bc_ackers decrementing to zero upon member departure does
    not restore *grp->open or trigger a socket wakeup) will be addressed
    in a separate patch.
    
    Fixes: 48a5fe38772b ("tipc: fix bc_ackers underflow on duplicate GRP_ACK_MSG")
    Fixes: 2f487712b893 ("tipc: guarantee that group broadcast doesn't bypass group unicast")
    Reported-by: James Burton <[email protected]>
    Cc: [email protected]
    Signed-off-by: Eric Dumazet <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
udp: relocate a connected socket in the 4-tuple hash table on re-connect [+ + +]
Author: Shardul Bankar <[email protected]>
Date:   Thu Sep 17 14:55:32 2026 +0530

    udp: relocate a connected socket in the 4-tuple hash table on re-connect
    
    [ Upstream commit 5fd0783b99d4af98f65cd58b56ec203d1d426104 ]
    
    A connected UDP socket that connects again to a different peer is not
    re-filed in the 4-tuple hash table:
    
        sk binds to 127.0.0.1:21001
        sk connects to 127.0.0.2:20001      // filed under hash(sk, peer1)
        sk connects to 127.0.0.3:20002      // still filed under hash(sk, peer1)
        packet from 127.0.0.3:20002         // hash(sk, peer2) misses, so the
                                            // lookup falls back to scoring the
                                            // hash2 chain for this address
                                            // and port
    
    udp_lib_hash4() returns early when the socket is already hashed, assuming
    ->rehash() relocates it. ->rehash() runs from __ip{4,6}_datagram_connect()
    only while the receive address is unset, which a second connect never is:
    the first connect assigns it, whether the socket was bound to a specific
    address or to the wildcard. commit 644f9108f3a5 ("udp: Make rehash4
    independent in udp_lib_rehash()") added that early return and named
    connect(AF_UNSPEC) as the way around it. That workaround does not help a
    socket with both SOCK_BINDADDR_LOCK and SOCK_BINDPORT_LOCK set, because
    __udp_disconnect() skips ->rehash() for the first and ->unhash() for the
    second.
    
    Delivery is correct either way.
    
    Relocate the socket when the hash it is filed under differs from the one
    requested, which is what commit 78c91ae2c6de ("ipv4/udp: Add 4-tuple hash
    for connected socket") did before the early return became unconditional. It
    is done here under hslot->lock, which that version did not take, to match
    udp_lib_rehash() and udp_lib_unhash(). hslot2 is unchanged, so hash4_cnt
    needs no adjustment, as in udp_lib_rehash(). A first connect is unaffected,
    and IPv6 shares the code.
    
    With 500 sockets on the port, a re-connected socket measured 522,553 pps
    without this change and 2,055,078 with it. The UDP side was noted as
    remaining work in [1].
    
    Link: https://lore.kernel.org/netdev/apnHqmYZQ4yzOP4N@v4bel/ [1]
    Fixes: 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()")
    Assisted-by: LLM
    Signed-off-by: Shardul Bankar <[email protected]>
    Reviewed-by: Kuniyuki Iwashima <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

udp: remove a disconnected socket from the 4-tuple hash table [+ + +]
Author: Shardul Bankar <[email protected]>
Date:   Thu Sep 17 14:55:33 2026 +0530

    udp: remove a disconnected socket from the 4-tuple hash table
    
    [ Upstream commit 9e95b1a94c9c49b4ba722251bbca2759b9c51737 ]
    
    A UDP socket bound to a specific address and port keeps its entry in the
    4-tuple hash table after it is disconnected:
    
        sk binds to 127.0.0.1:21001
        sk connects to 127.0.0.2:20001      // filed in the 4-tuple table
        sk disconnects, connect(AF_UNSPEC)  // still filed, peer now 0.0.0.0:0
    
    __udp_disconnect() takes a socket out of that table only as a side effect
    of ->rehash() or ->unhash(), and it skips ->rehash() when
    SOCK_BINDADDR_LOCK is set and ->unhash() when SOCK_BINDPORT_LOCK is set.
    commit 6996a2d2d0a6 ("udp: Unhash auto-bound connected sk from 4-tuple hash
    table when disconnected.") fixed the same end state for a wildcard-bound
    socket, by a path this one does not take.
    
    The entry is counted whether or not anything hits it. hash4_cnt on the
    hash2 slot stays raised for as long as the socket lives, so udp_has_hash4()
    keeps sending every packet for that address and port through the 4-tuple
    lookup first.
    
    On IPv6 it can also be hit. __udp_disconnect() does not clear sk_v6_daddr,
    so udp_v6_rehash() files the entry under the peer the socket was connected
    to with a zero dport, and inet6_match() compares that same
    field: a datagram from the former peer with a zero source port matches,
    and source port zero is accepted on receive. On IPv4 the peer is cleared,
    so a match would need a zero source address as well, which the routing
    layer rejects as martian. The stale sk_v6_daddr is a separate defect, not
    addressed here; removing the entry closes this path either way.
    
    The entry can also be relocated. __udp_disconnect() clears sk_bound_dev_if,
    so a subsequent SO_BINDTODEVICE calls ->rehash(), and because the receive
    address is still specific udp_lib_rehash() moves the entry instead of
    removing it, into the bucket that (rcv_saddr, num, 0, 0) hashes to -- a
    pure function of the address and port, so every socket reaching this state
    on one address and port collects in one bucket. The bucket cannot be chosen
    from outside, as udp_ehashfn() is seeded with a per-boot secret. This last
    one became reachable only with commit 644f9108f3a5 ("udp: Make rehash4
    independent in udp_lib_rehash()"), which moved the hash4 handling out of a
    branch a disconnected socket does not take; the stale entry itself dates
    from the commit in Fixes.
    
    Take the socket out of the table before __udp_disconnect() runs, while it
    still matches how it was filed. This also reaches the wildcard case ahead
    of udp_lib_rehash()'s udp_unhash4() branch, leaving that branch unreachable
    from udp_disconnect(); removing it belongs in net-next. udp_disconnect()
    and udp_abort() are the only UDP entries into __udp_disconnect(), which is
    shared with raw, ping and l2tp sockets that are not struct udp_sock:
    ping_prot.obj_size is sizeof(struct inet_sock), so udp_hashed4() on one
    would read past the allocation.
    
    Fixes: 78c91ae2c6de ("ipv4/udp: Add 4-tuple hash for connected socket")
    Assisted-by: LLM
    Signed-off-by: Shardul Bankar <[email protected]>
    Reviewed-by: Kuniyuki Iwashima <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Paolo Abeni <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
veth: manage XDP program pointers during channel resize [+ + +]
Author: Jakub Kicinski <[email protected]>
Date:   Mon Sep 21 16:18:56 2026 -0700

    veth: manage XDP program pointers during channel resize
    
    [ Upstream commit 7104a370714346b667712913dc16abf14bbc97ed ]
    
    veth_set_channels() tears down XDP resources for removed RX queues
    without clearing rq->xdp_prog.  If the program is then detached or
    replaced, those queues keep the old pointer after bpf_prog_put().
    A later channel increase can re-enable NAPI and run the freed program.
    
      BUG: unable to handle page fault for address: ffffc90000256048
      Oops: Oops: 0000 [#1] SMP KASAN NOPTI
      RIP: veth_xdp_rcv_skb (include/linux/filter.h:779
                             include/net/xdp.h:696 drivers/net/veth.c:820)
      Call Trace:
       veth_xdp_rcv (drivers/net/veth.c:941)
       veth_poll (drivers/net/veth.c:986)
       __napi_poll (net/core/dev.c:7787)
       net_rx_action (net/core/dev.c:7850 net/core/dev.c:8007)
       handle_softirqs (kernel/softirq.c:645)
      Kernel panic - not syncing: Fatal exception in interrupt
    
    Fixes: 4752eeb3d891 ("veth: implement support for set_channel ethtool op")
    Signed-off-by: Weiming Shi <[email protected]>
    Acked-by: Stanislav Fomichev <[email protected]>
    Reviewed-by: Jiayuan Chen <[email protected]>
    Reviewed-by: Jason Xing <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
virtio_net: copy zerocopy frags in start_xmit without NAPI [+ + +]
Author: Willem de Bruijn <[email protected]>
Date:   Fri Sep 18 20:47:28 2026 -0400

    virtio_net: copy zerocopy frags in start_xmit without NAPI
    
    commit 07e1a9408b6c2f9d0cfb757b67dabb52da7a32b2 upstream.
    
    Virtio-net without NAPI frees completed skbs lazily on the next
    start_xmit. Senders waiting for in-flight zerocopy buffers can
    deadlock if they cannot transmit more packets, as then no
    completed packets will be freed.
    
    When !use_napi, virtio-net already calls skb_orphan to avoid waiting
    up for transmitted skbs to be freed. For zerocopy packets that
    require deep copying on orphan (i.e. those that do not set
    SKBFL_DONT_ORPHAN, such as PACKET_TX_RING), call skb_orphan_frags
    before orphaning to release the buffers.
    
    This fixes the tpacket_snd slot reuse bug on skb_orphan for
    virtio-net, and prevents PACKET_TX_RING from running out of slots.
    
    This fix also touches vhost_net zerocopy packets, which also do not
    set SKBFL_DONT_ORPHAN. This is fine: vhost_net packets only encounter
    virtio-net in nested virtualization, and only if napi_tx is
    explicitly disabled (it has been default-enabled since Linux 4.12).
    In that rare case, copying the frags is desirable anyway to prevent
    holding guest descriptors pinned across unbounded intervals.
    
    This is a prerequisite for the next patch, which converts
    PACKET_TX_RING to standard zerocopy completion. Without this patch
    first, a bounded ring sender can stall indefinitely behind a
    virtio-net virtqueue that cannot reclaim.
    
    Fixes: 5cd8d46ea156 ("packet: copy user buffers before orphan or clone")
    Cc: [email protected]
    Cc: [email protected]
    Cc: [email protected]
    Signed-off-by: Willem de Bruijn <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
vlan: ensure sufficient headroom in vlan_dev_hard_header() [+ + +]
Author: Eric Dumazet <[email protected]>
Date:   Thu Sep 24 08:29:51 2026 +0000

    vlan: ensure sufficient headroom in vlan_dev_hard_header()
    
    [ Upstream commit cd5dd68267c4238795fadaf02b3575ca3f8a6500 ]
    
    Callers that only reserve ETH_HLEN or less (such as llc_alloc_frame()),
    or skbs allocated before dynamic device/headroom changes (e.g. toggling
    VLAN_FLAG_REORDER_HDR or bonding/team switching slaves), can reach
    vlan_dev_hard_header() with insufficient headroom and trigger
    skb_under_panic().
    
    Use skb_cow_head() in vlan_dev_hard_header() when VLAN_FLAG_REORDER_HDR
    is not set to ensure sufficient headroom for the VLAN header(s) and the
    underlying device hard header.
    
    Use READ_ONCE() to read dev->hard_header_len and dev->needed_headroom as
    they can be updated concurrently under RTNL (e.g. in
    vlan_transfer_features()) while vlan_dev_hard_header() runs locklessly on
    the transmit path. Also avoid LL_RESERVED_SPACE(dev) here so that the
    extra HH_DATA_MOD alignment padding does not trigger unnecessary
    pskb_expand_head() reallocations on inner stacked VLAN devices after the
    outer VLAN header has been pushed.
    
    Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
    Reported-by: Zixuan Chai <[email protected]>
    Closes: https://lore.kernel.org/netdev/[email protected]/
    Link: https://lore.kernel.org/netdev/[email protected]/
    Cc: Hangbin Liu <[email protected]>
    Signed-off-by: Eric Dumazet <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

vlan: require the MAC header to be present in __vlan_insert_inner_tag() [+ + +]
Author: Xiang Mei <[email protected]>
Date:   Tue Sep 15 01:31:52 2026 -0700

    vlan: require the MAC header to be present in __vlan_insert_inner_tag()
    
    commit ab888242fce4f16f6c4d4c6ec53939ad36aa3b3a upstream.
    
    __vlan_insert_inner_tag() only guarantees head room via skb_cow_head(),
    never that mac_len bytes of MAC header are present.  Its ETH_HLEN
    wrappers - __vlan_insert_tag() under skb_vlan_push(), and
    vlan_insert_tag() under validate_xmit_vlan() on the generic transmit
    path - therefore rewrite the first 16 bytes at skb->data: a 12-byte
    memmove plus two 2-byte stores at +12 and +14.  No caller supplies the
    bound, while the pop helpers use skb_ensure_writable()/pskb_may_pull().
    
    An IFF_TUN device has hard_header_len == 0, so packet_snd() accepts a
    one-byte AF_PACKET/SOCK_RAW frame.  The first vlan push only sets a
    hwaccel tag; the next - clsact "action vlan push" or
    bpf_skb_vlan_push() - enters the helper with skb->len still 1.  The
    head comes from skbuff_small_head without __GFP_ZERO, so each push
    drags bytes from beyond skb->tail into the frame.  After three the
    one-byte send leaves as 13 bytes carrying 11 bytes of uninitialised
    slab:
    
      0000: 5a b3 62 12 80 88 ff ff 00 b3 62 12 81
               `------------------------------'
      only 0x5a was sent; the rest is slab, here the top 56 bits of a
      linear-map address
    
    Require the MAC header the helper rewrites to be present, so such a
    frame is dropped rather than transmitted.
    
    Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
    Cc: [email protected]
    Reported-by: [email protected]
    Signed-off-by: Xiang Mei <[email protected]>
    Reviewed-by: Simon Horman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
vrf: Stop corrupting skb->csum when capturing CHECKSUM_COMPLETE packets [+ + +]
Author: Ido Schimmel <[email protected]>
Date:   Tue Sep 22 16:12:39 2026 +0300

    vrf: Stop corrupting skb->csum when capturing CHECKSUM_COMPLETE packets
    
    [ Upstream commit ab7aa05c06ae340e5c7530bb78fa8d23794e460b ]
    
    The VRF device is an Ethernet device but it can have non-Ethernet ports
    such as IP tunnels. Before the cited commit, capturing packets from such
    ports on the VRF device resulted in these packets being detected as
    malformed since they lack an Ethernet header.
    
    The cited commit fixed it by pushing a dummy Ethernet header to such
    packets before the capture and pulling it afterwards. In the case of
    CHECKSUM_COMPLETE packets it also updated skb->csum with the checksum of
    the dummy Ethernet header. This is wrong as skb->csum should not include
    the checksum of the Ethernet header ("checksum of the _whole_ packet as
    seen by netif_rx()").
    
    This also means that L4 protocols receive a corrupted skb->csum and
    potentially drop the packet, as is the case with UDP packets whose
    checksum was completed by software.
    
    Fix by removing the unnecessary call to skb_postpush_rcsum().
    
    Fixes: 048939088220 ("vrf: add mac header for tunneled packets when sniffer is attached")
    Reported-by: Stefano Sasso <[email protected]>
    Closes: https://lore.kernel.org/netdev/CALtE316UtL3x7LL6uxfXzx8rW6AbzYPeDOb478hqJCr_-dj=Wg@mail.gmail.com/
    Signed-off-by: Ido Schimmel <[email protected]>
    Reviewed-by: David Ahern <[email protected]>
    Reviewed-by: Eric Dumazet <[email protected]>
    Reviewed-by: Andrea Mayer <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
vsock: ignore empty child namespace mode writes [+ + +]
Author: Aldo Ariel Panzardo <[email protected]>
Date:   Tue Sep 15 14:30:50 2026 -0300

    vsock: ignore empty child namespace mode writes
    
    commit 2ec28c09b320ba241bea8a70ee5cb9ccf4a099e8 upstream.
    
    __vsock_net_mode_string() returns success without updating new_mode when
    the transfer length is zero. Its caller then reads the uninitialized enum
    and may permanently store a stack-derived value in the write-once child
    mode.
    
    Return before calling __vsock_net_mode_string() when *lenp is zero so
    that the helper is never invoked with nothing to parse and new_mode is
    never read uninitialized. This also prevents an empty write from
    locking the current mode.
    
    Fixes: eafb64f40ca4 ("vsock: add netns to vsock core")
    Cc: [email protected]
    Reviewed-by: Luigi Leonardi <[email protected]>
    Signed-off-by: Aldo Ariel Panzardo <[email protected]>
    Reviewed-by: Stefano Garzarella <[email protected]>
    Reviewed-by: Bobby Eshleman <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
vxlan: use one headroom snapshot for neighbour replies [+ + +]
Author: Sanghyun Park <[email protected]>
Date:   Fri Sep 18 12:26:58 2026 +0900

    vxlan: use one headroom snapshot for neighbour replies
    
    [ Upstream commit 481506a756dcd828ef42391cb08f38d8d96d38fc ]
    
    vxlan_na_create() samples LL_RESERVED_SPACE() to size the reply skb and then
    samples it again to reserve headroom. A concurrent vxlan_changelink() can
    update needed_headroom between the two reads, creating a TOCTOU race. The
    second value can exceed the allocation and make the Ethernet header write out
    of bounds.
    
    The race is reproducible on the unpatched kernel. It occurred when
    vxlan_na_create() generated a neighbour reply while vxlan_changelink() changed
    the link headroom. KASAN caught a four-byte write two bytes beyond a 704-byte
    skbuff_small_head allocation.
    
    Snapshot the headroom once and use that value for both allocation and
    reservation.
    
    Fixes: 4b29dba9c085 ("vxlan: fix nonfunctional neigh_reduce()")
    Signed-off-by: Sanghyun Park <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Jakub Kicinski <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>

 
workqueue: Fix NULL current_pwq deref in flush dependency check [+ + +]
Author: Pavankumar Kondeti <[email protected]>
Date:   Fri Sep 25 15:25:12 2026 +0530

    workqueue: Fix NULL current_pwq deref in flush dependency check
    
    commit db6365ced4d5855e321f772b240c0e473bcfcdd5 upstream.
    
    check_flush_dependency() uses current_wq_worker() to determine whether
    the caller is a workqueue worker and then dereferences worker->current_pwq
    to test whether the current workqueue is WQ_MEM_RECLAIM.
    
    current_wq_worker() only means that %current has PF_WQ_WORKER set. A
    kworker can reach check_flush_dependency() while it is not executing a
    work item. One such path is worker_thread() acting as the pool manager,
    where create_worker() does GFP_KERNEL allocation and the allocation path
    invokes the OOM notifier. In that state worker->current_pwq is NULL
    because current_pwq is set only by process_one_work() and cleared again
    after the work function returns.
    
    [  416.760634][  T375] Call trace:
    [  416.760638][  T375]  check_flush_dependency+0x80/0x120 (P)
    [  416.760648][  T375]  __flush_work+0x98/0x224
    [  416.760657][  T375]  flush_work+0x30/0x44
    [  416.760665][  T375]  ...
    [  416.760710][  T375]  blocking_notifier_call_chain+0x58/0xa0
    [  416.760719][  T375]  out_of_memory+0xb4/0x458
    [  416.760730][  T375]  __alloc_pages_may_oom+0x11c/0x1a8
    [  416.760739][  T375]  __alloc_pages_slowpath+0x314/0x46c
    [  416.760746][  T375]  __alloc_frozen_pages_noprof+0x110/0x1a4
    [  416.760753][  T375]  new_slab+0x12c/0x484
    [  416.760759][  T375]  ___slab_alloc+0x7a8/0xc7c
    [  416.760765][  T375]  __slab_alloc+0x74/0xd8
    [  416.760772][  T375]  __kmalloc_cache_node_noprof+0x2ac/0x304
    [  416.760779][  T375]  alloc_worker+0x28/0x60
    [  416.760785][  T375]  create_worker+0x4c/0x20c
    [  416.760790][  T375]  worker_thread+0xe8/0x2b8
    [  416.760796][  T375]  kthread+0x1a8/0x200
    [  416.760805][  T375]  ret_from_fork+0x10/0x20
    
    Guard the WQ_MEM_RECLAIM-worker warning with worker->current_pwq. If the
    kworker is not currently executing a work item, there is no current
    workqueue to diagnose with that warning. The PF_MEMALLOC warning is left
    unchanged so explicit reclaim context flushing a !WQ_MEM_RECLAIM target
    is still reported.
    
    Fixes: fca839c00a12 ("workqueue: warn if memory reclaim tries to flush !WQ_MEM_RECLAIM workqueue")
    Cc: [email protected]
    Assisted-by: LLM
    Signed-off-by: Pavankumar Kondeti <[email protected]>
    Signed-off-by: Tejun Heo <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
writeback: bound cleanup_offline_cgwb() rescans by rotating scanned inodes [+ + +]
Author: Patrick Lu (Anthropic) <[email protected]>
Date:   Fri Sep 11 18:49:49 2026 +0000

    writeback: bound cleanup_offline_cgwb() rescans by rotating scanned inodes
    
    commit f6988c90671e83db79df1b7b9d6fdb0e5947fd84 upstream.
    
    cleanup_offline_cgwb() prepares at most WB_MAX_INODES_PER_ISW inodes
    per call and is called again until the dying wb is drained, but every
    call walks wb->b_attached and then wb->b_dirty_time from the same end.
    Inodes already prepared (they stay on the list with I_WB_SWITCH set
    until the switch worker runs) and inodes that cannot be switched
    (I_FREEING, I_WILL_FREE, !SB_ACTIVE, DAX, already on the target wb)
    stay where they are, so each pass rescans a growing run of them under
    wb->list_lock and a full drain is quadratic in the number of inodes on
    the list. With ~17M inodes attached to one dying cgwb we saw this end
    in soft lockups, with CPUs reported stuck for 21-48s.
    
    Walk both lists from the oldest end and move every scanned inode to
    the newest end, so the next pass starts where the previous one stopped
    and the drain becomes linear. b_attached is unordered, so nobody sees
    the reorder there. b_dirty_time is ordered by dirtied_when, but the
    oldest unscanned inode stays at the end move_expired_inodes() picks
    from, sync takes the whole list regardless of order, and prepared
    inodes leave the list as soon as the switch work runs and get a new
    dirtied_time_when on the new wb anyway, so the only inodes left out of
    order are the ones that can never switch (DAX), and only on the dying
    wb.
    
    Fixes: c22d70a162d3 ("writeback, cgroup: release dying cgwbs by switching attached inodes")
    Cc: [email protected]
    Acked-by: Tejun Heo <[email protected]>
    Acked-by: Roman Gushchin <[email protected]>
    Signed-off-by: Patrick Lu (Anthropic) <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Reviewed-by: Jan Kara <[email protected]>
    Signed-off-by: Christian Brauner (Amutable) <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

writeback: report a Tasks-RCU quiescent state per cgwb drain pass [+ + +]
Author: Josef Bacik <[email protected]>
Date:   Wed Sep 9 18:01:07 2026 +0000

    writeback: report a Tasks-RCU quiescent state per cgwb drain pass
    
    commit 407a5d205179a4ab186571b0e16ec42725dc77bc upstream.
    
    cleanup_offline_cgwbs_workfn() drains a dying cgwb by calling
    cleanup_offline_cgwb() until it returns false, with a cond_resched()
    between passes.  On a CONFIG_PREEMPTION kernel that cond_resched() does
    nothing: _cond_resched() is a plain "return 0", and under PREEMPT_DYNAMIC
    the full and lazy modes disable it.  Since commit 7dadeaa6e851 ("sched:
    Further restrict the preemption modes") those are the only two models on
    the architectures with PREEMPT_LAZY support, arm64 and x86 among them, so
    the drain loop never reports a Tasks-RCU quiescent state.
    
    A worker draining a cgwb with millions of attached inodes runs for
    minutes.  On a 6.18 arm64 host in lazy mode the cgwb worker drained one
    dying cgroup's writeback domain for over 11 minutes.  A BPF program unlink
    (bpf_trampoline_unlink_prog -> bpf_trampoline_update ->
    unregister_ftrace_direct -> ftrace_shutdown -> synchronize_rcu_tasks())
    waited on that grace period while holding the trampoline mutex, 42 tasks
    queued behind it in D state, and the hung task detector fired at 614 s and
    panicked the host.  Any BPF or ftrace detach during a long drain inherits
    the drain's length.
    
    Fix this by calling cond_resched_tasks_rcu_qs() so we do not stall out
    anybody who calls sycnrhonize_rcu_tasks().  We put this in a do { } while
    loop because if we have many small cgroups cleanup_offline_cgwb() will
    return false and we will never call cond_resched_tasks_rcu_qs(), creating
    the same problem.
    
    Link: https://lore.kernel.org/[email protected]
    Fixes: c22d70a162d3 ("writeback, cgroup: release dying cgwbs by switching attached inodes")
    Signed-off-by: Josef Bacik <[email protected]>
    Signed-off-by: Andrew Morton <[email protected]>
    Link: https://lore.kernel.org/bpf/9d444098-7c03-4163-af12-bd0a79a51443@paulmck-laptop/
    Assisted-by: LLM
    Acked-by: Tejun Heo <[email protected]>
    Reviewed-by: Roman Gushchin <[email protected]>
    Reviewed-by: Jan Kara <[email protected]>
    Acked-by: Lorenzo Stoakes (ARM) <[email protected]>
    Cc: David Hildenbrand <[email protected]>
    Cc: Dennis Zhou <[email protected]>
    Cc: Liam R. Howlett <[email protected]>
    Cc: Matthew Wilcox (Oracle) <[email protected]>
    Cc: Michal Hocko <[email protected]>
    Cc: Mike Rapoport <[email protected]>
    Cc: "Paul E . McKenney" <[email protected]>
    Cc: Suren Baghdasaryan <[email protected]>
    Cc: Vlastimil Babka <[email protected]>
    Cc: <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
x86/mce: Fix hardware debug register corruption on task migration [+ + +]
Author: Masami Hiramatsu (Google) <[email protected]>
Date:   Tue Sep 22 13:24:55 2026 +0900

    x86/mce: Fix hardware debug register corruption on task migration
    
    commit b8d1d5b63a8ef532038eebd9d97d406860385668 upstream.
    
    In exc_machine_check_user(), local_db_save() and local_db_restore() are
    invoked in the outer entry stubs (DEFINE_IDTENTRY_MCE_USER,
    DEFINE_FREDENTRY_MCE, and DEFINE_IDTENTRY_RAW), surrounding
    exc_machine_check_user().
    
    However, exc_machine_check_user() calls irqentry_exit_to_user_mode(), which
    handles pending thread work and may schedule() if TIF_NEED_RESCHED is set. If
    the task migrates to another CPU during schedule(), local_db_restore() runs on
    the new CPU with the dr7 state saved from the old CPU. This corrupts the new
    CPU's DR7 hardware debug register and leaves the old CPU's DR7 disabled.  In
    short, local_db_save() and local_db_restore() pair must be run on the same
    CPU.
    
    To fix this, move local_db_save() and local_db_restore() inside
    exc_machine_check_user() and exc_machine_check_kernel(). In
    exc_machine_check_user(), DR7 is saved and restored strictly around
    do_machine_check() to avoid schedule() during migration. In
    exc_machine_check_kernel(), local_db_save() is called at the entry point to
    prevent early memory accesses from triggering nested #DB exceptions, and
    restored on all exits.
    
    Fixes: cd840e424f27 ("x86/entry, mce: Disallow #DB during #MC")
    Assisted-by: LLM
    Signed-off-by: Masami Hiramatsu (Google) <[email protected]>
    Signed-off-by: Borislav Petkov (AMD) <[email protected]>
    Acked-by: Peter Zijlstra (Intel) <[email protected]>
    Cc: <[email protected]>
    Link: https://patch.msgid.link/179005109564.388919.3937970081044095776.stgit@devnote2
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
x86/PCI: Disable enhanced atomics on AMD NBIO 7.7 and 7.11 [+ + +]
Author: Mario Limonciello <[email protected]>
Date:   Tue Sep 8 14:05:59 2026 -0500

    x86/PCI: Disable enhanced atomics on AMD NBIO 7.7 and 7.11
    
    commit 4fde448225123442c5796f54b7a4400e2d3cbaf6 upstream.
    
    Multiple users report data corruption during 64-bit DMA transfers on
    systems with AMD NBIO 7.7 and 7.11 controllers.
    
    This occurs when BIOS enables AMD "enhanced atomic operations" on PCIe Root
    Ports. When enhanced atomics are enabled, any 64-bit DMA access may be
    corrupted.
    
    Disable enhanced atomics using SMN for NBIO 7.7 and 7.11 based models.
    
    Reported-by: Mikael Etienne <[email protected]>
    Closes: https://lore.kernel.org/[email protected]/
    Reported-by: Arthur Husband <[email protected]>
    Closes: https://lore.kernel.org/[email protected]/
    Reported-by: Alvin Lim <[email protected]>
    Closes: https://lore.kernel.org/[email protected]/
    Signed-off-by: Mario Limonciello <[email protected]>
    [bhelgaas: commit log, s/IOVA/DMA/ in comment]
    Signed-off-by: Bjorn Helgaas <[email protected]>
    Cc: [email protected]
    Cc: David Laight <[email protected]>
    Cc: John Smith <[email protected]>
    Cc: Lennert Buytenhek <[email protected]>
    Cc: Niklas Cassel <[email protected]>
    Cc: Roland Waltersson <[email protected]>
    Link: https://patch.msgid.link/[email protected]
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
x86/sev: Make vTPM SVSM calls preemption-safe [+ + +]
Author: Melody Wang <[email protected]>
Date:   Mon Sep 14 00:57:17 2026 +0000

    x86/sev: Make vTPM SVSM calls preemption-safe
    
    commit 6c43c72748fffd29dec15cd1f31e9a32949bc437 upstream.
    
    Two functions in the SVSM vTPM guest implementation do not disable
    preemption when fetching the SVSM Calling Area Address (CAA).
    
    The SVSM CAA is a per-CPU structure. When a thread is preempted and migrated
    to a different CPU after fetching the per-CPU CAA, the SVSM call will execute
    on the new CPU with the original CPU's CAA. Which is wrong.
    
    Move the CAA fetching operation inside svsm_perform_call_protocol() which
    disables interrupts around the SVSM call and thus runs preemption-safe.
    
    Fixes: 770de678bc28 ("x86/sev: Add SVSM vTPM probe/send_command functions")
    Signed-off-by: Melody Wang <[email protected]>
    Signed-off-by: Borislav Petkov (AMD) <[email protected]>
    Reviewed-by: Stefano Garzarella <[email protected]>
    Cc: [email protected]
    Link: https://patch.msgid.link/a5bc0d4a2c462a0089109e145c21626b244b2ff0.1789345277.git.huibo.wang@amd.com
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
xfs: call xfs_dquot_set_prealloc_limits if we installed default rtb limits [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Mon Sep 14 22:37:05 2026 -0700

    xfs: call xfs_dquot_set_prealloc_limits if we installed default rtb limits
    
    commit 065f3ce5936e68da75f3dc18201d290073d78f3c upstream.
    
    Now that we have quotas for the realtime volume, we also have
    precomputed watermark limits for the realtime block counts.  These
    precomputations should be done any time we change the rtb limits, which
    means that xfs_qm_adjust_dqlimits needs to ensure that if we installed
    a default rtb limit.
    
    Cc: [email protected] # v6.13
    Fixes: 5dd70852b03901 ("xfs: create quota preallocation watermarks for realtime quota")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: check di_forkoff correctly in scrub [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Thu Sep 10 22:54:58 2026 -0700

    xfs: check di_forkoff correctly in scrub
    
    commit e9193f2f1ce32d02b9230094ffdbfab715ab6137 upstream.
    
    The di_forkoff check in xchk_dinode is incorrect, according to LOLLM.
    XFS_DFORK_BOFF returns a byte count relative to the start of the literal
    area, not the start of the inode.  Therefore, this check won't flag
    di_forkoff values that are larger than the literal area but not the
    inode size itself.  Fix this check; sadly the old APTR code was correct.
    
    Cc: [email protected] # v6.8
    Fixes: 6b5d917780219d ("xfs: dont cast to char * for XFS_DFORK_*PTR macros")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: check padding field in xfs_ioc_commit_range [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Thu Sep 10 22:53:39 2026 -0700

    xfs: check padding field in xfs_ioc_commit_range
    
    commit 3083ba8dde765a9ab2337f3db68d00724a6b1202 upstream.
    
    LOLLM points out that we don't check the ioctl padding field here, so
    let's do that.  I don't think there are many users yet since exchrange
    requires a new feature flag, so it's a good time to try to plug this
    hole.
    
    Cc: [email protected] # v6.12
    Fixes: 398597c3ef7fb1 ("xfs: introduce new file range commit ioctls")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: don't assert when XFS_SCRUB_TYPE_HEALTHY scans return corruption [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Wed Sep 9 23:00:16 2026 -0700

    xfs: don't assert when XFS_SCRUB_TYPE_HEALTHY scans return corruption
    
    commit afbccf99f7f82117cba9ad4b0b006692030f49e8 upstream.
    
    XFS_SCRUB_TYPE_HEALTHY is a synthentic scrub type so that xfs_scrub can
    tell the kernel "Hey, I finished a scan and saw no problems" and have
    the kernel forget that it saw indirect evidence of corruption.
    
    Unfortunately, as LOLLM points out, it's possible for the health system
    to record a new corruption just before xfs_scrub gets to
    XFS_SCRUB_TYPE_HEALTHY.  In this case, the existing logic doesn't return
    early and instead wanders into unknown regions of type_to_health_flag
    and trips the assert because HEALTHY doesn't have a group assignment.
    
    Fix the logic so that we always return early for a HEALTHY scrub type,
    even if we decide not to call xchk_mark_all_healthy.
    
    Cc: [email protected] # v6.9
    Fixes: a1f3e0cca41036 ("xfs: update health status if we get a clean bill of health")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: don't call xfs_exchange_range_finish for a dry run [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Thu Sep 10 22:53:55 2026 -0700

    xfs: don't call xfs_exchange_range_finish for a dry run
    
    commit 8fc18580ec17f90beac4c933fbe4c74dcd3b7f36 upstream.
    
    LOLLM noticed that we strip file privileges and whatnot even for a dry
    run.  We also shouldn't flush dirty data to disk or trim COW staging
    events for a dry run.  Neither of those behaviors are allowed by the
    manpage, so fix that by exiting early on DRY_RUN in various functions.
    
    Cc: [email protected] # v6.10
    Fixes: 42672471f938cd ("xfs: bind together the front and back ends of the file range exchange code")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: don't let hidden_space go negative in xfs_metafile_resv_init [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Mon Sep 14 22:39:10 2026 -0700

    xfs: don't let hidden_space go negative in xfs_metafile_resv_init
    
    commit 476582d754cdc5110f806001417fea6c77824c13 upstream.
    
    LOLLM points out that if the amount of fdblocks that we can reserve for
    a metadata btree file goes below the space already used by that file,
    then the hidden_space subtraction can underflow, causing
    xfs_dec_fdblocks to subtract a huge amount of space.  We never want the
    target to be less than the used sapce, so fix the logic that adjusts
    dblocks_avail downwards.
    
    Also fix an error in the adjacent comment.
    
    Cc: [email protected] # v6.15
    Fixes: 1df8d75030b787 ("xfs: make metabtree reservations global")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: don't let memory failures leak blocks and kill repairs [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Mon Sep 14 22:37:52 2026 -0700

    xfs: don't let memory failures leak blocks and kill repairs
    
    commit ab1c416d2377cdc16123ef521aac4da1c468c3d4 upstream.
    
    LOLLM complains that a memory allocation failure in
    xrep_newbt_add_blocks results in online repair leaking blocks that were
    previously allocated to write a new btree, but the problem is worse than
    that -- a limitation of the codebase is that the callers cannot undo the
    transaction /and/ return the error -- either you undo all changes and
    commit the transaction, or you error out and the filesystem goes down.
    
    However, the new btree space reservation object isn't that big (~48
    bytes).  Let's just do a NOFAIL allocation and the problem goes away.
    
    Cc: [email protected] # v6.8
    Fixes: be408417630427 ("xfs: implement block reservation accounting for btrees we're staging")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: don't merge different file IO error types [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Mon Sep 14 22:38:07 2026 -0700

    xfs: don't merge different file IO error types
    
    commit d80993655f7be2a461c0a63e2022da5757f47dac upstream.
    
    LOLLM noticed that we can accidentally merge file range health
    monitoring events even if they have different errors.  We shouldn't do
    that.
    
    Cc: [email protected] # v7.0
    Fixes: dfa8bad3a8796c ("xfs: convey file I/O errors to the health monitor")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: drop dquot flush lock when we can't find a buffer to flush [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Mon Sep 14 22:38:54 2026 -0700

    xfs: drop dquot flush lock when we can't find a buffer to flush
    
    commit ffb48dccce1960a9ea24463a2f3c21d124d6b672 upstream.
    
    LOLLM noticed that xfs_qm_flush_one fails to drop the dquot flush lock
    if it can't grab the buffer associated with the dquot.  Since there's no
    buffer, nobody else is going to drop the dqflock, so we need to do it
    ourselves.
    
    Cc: [email protected] # v6.13
    Fixes: ca378189fdfa89 ("xfs: convert quotacheck to attach dquot buffers")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: fix attr fork block count checks in xrep_inode_blockcounts [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Wed Sep 9 23:00:31 2026 -0700

    xfs: fix attr fork block count checks in xrep_inode_blockcounts
    
    commit bb991b7f79dd34cc5f24db0f736bf75630c970e7 upstream.
    
    LOLLM points out that a file has an attr fork, it will call
    xchk_inode_count_blocks to set @ablocks to the number of fsblocks mapped
    by the attr fork; but then it'll compare @blocks (aka the count of
    fsblocks mapped by the data fork).  We already checked that and we never
    do anything with @acount, so I think this is clearly a bug.  Fix the
    comparison.
    
    Cc: [email protected] # v6.8
    Fixes: 2d295fe65776d1 ("xfs: repair inode records")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: fix blockgc group quota scanning when usrquota isn't enforced [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Mon Sep 14 22:38:23 2026 -0700

    xfs: fix blockgc group quota scanning when usrquota isn't enforced
    
    commit f8f6382ff13109e19d0fc1d0224ff7d29641e56a upstream.
    
    LOLLM noticed the copy-paste error here -- if user quotas aren't
    enforced but we're near the group quota limit, we fail to set FLAG_GID
    and hence we might not actually free any preallocations, causing
    unnecessary EDQUOT.  Fix that.
    
    Cc: [email protected] # v5.12
    Fixes: c237dd7c709432 ("xfs: flush eof/cowblocks if we can't reserve quota for inode creation")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: fix cursor and pointer handling when recovering iunlink buckets [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Mon Sep 14 22:38:38 2026 -0700

    xfs: fix cursor and pointer handling when recovering iunlink buckets
    
    commit 65f39d09d73718611cee40399179323b5d4ead00 upstream.
    
    LOLLM pointed out a bug in xlog_recover_iunlink_bucket:
    
    1. We don't null out prev_ip after releasing it, which can lead to UAF
       problems if the inodegc flush call in the loop fails.
    
    at which point I noticed even more bugs:
    
    2. If the inodegc flush inside the loop fails, we also leak @ip.
    
    3. We set prev_agino to agino having already advanced agino, which
       results in inodes with i_prev_unlinked set to itself.
    
    4. If we exit the bottom of the loop with prev_ip set, then prev_ip
       aliases ip and we also set its i_prev_unlinked to itself.
    
    Bugs 3 and 4 introduce loops into the unlinked list, though these loops
    don't surface because we immediately flush each unlinked inode after
    loading it.
    
    Fix all of these issues.
    
    Cc: [email protected] # v6.0
    Fixes: 04755d2e5821b3 ("xfs: refactor xlog_recover_process_iunlinks()")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: fix rtgroup repair estimations [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Mon Sep 14 22:37:20 2026 -0700

    xfs: fix rtgroup repair estimations
    
    commit 41c4c41cf6c44f98db2916e1781f537d9ba6461a upstream.
    
    When I added online fsck for realtime reflink, I forgot to update
    xrep_calc_rtgroup_resblks to factor in the size of the refcount btree
    when it guesses how much space we need to start a repair.  This hasn't
    been a huge problem in practice because there are few filesystems with
    (a) realtime, (b) rtgroups, (c) reflink, and (d) no rmap.  But let's fix
    this before someone stumbles upon it, especially since LOLLM flagged
    this for me.
    
    Cc: [email protected] # v6.14
    Fixes: 83ccffc489975d ("xfs: online repair of the realtime refcount btree")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: fix wild memcpy access when formatting ondisk rtrefcount btree roots [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Mon Sep 14 22:39:25 2026 -0700

    xfs: fix wild memcpy access when formatting ondisk rtrefcount btree roots
    
    commit fe2f9135df43db849e74f03956d322ac20b59af7 upstream.
    
    LOLLM noticed that the inode btree root formatting methods copy too many
    bytes -- there's only one set of keys in node blocks, not two.  This
    causes memory corruption of whatever's beyond the buffers.
    
    Cc: [email protected] # v6.14
    Fixes: f0415af60f482a ("xfs: wire up a new metafile type for the realtime refcount")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: only flag zero padding for dir3 data blocks, not dir3 block blocks [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Thu Sep 10 22:54:42 2026 -0700

    xfs: only flag zero padding for dir3 data blocks, not dir3 block blocks
    
    commit c54110d814c3e8ed6bbbb5e02874994ed9f668ac upstream.
    
    LOLLM complains that xchk_directory_data_bestfree can be passed a
    directory block that is either in "block" or "data" format, but the
    check here unconditionally treats the dir3_block and dir3_data blocks as
    if they have the same header format (they don't).  Consequently, we can
    incorrectly set the preen state on dir3_block blocks, which of course
    we can't preen away because dir3_block blocks do not have a padding
    field.  Fix this.
    
    Cc: [email protected] # v7.1-rc4
    Fixes: 939919ccddfcc3 ("xfs: check directory data block header padding in scrub")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: release orphanage dir inode if chown fails [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Wed Sep 9 23:00:47 2026 -0700

    xfs: release orphanage dir inode if chown fails
    
    commit 1c32cdc986467eaffeedb6c5334852809555b82d upstream.
    
    LOLLM points out that we leak the igrab'd reference to the orphanage
    directory inode if chowning it fails.  Fix that.
    
    Cc: [email protected] # v6.10
    Fixes: 1e58a8ccf2597c ("xfs: move orphan files to the orphanage")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: use correct jiffies comparison function in xchk_maybe_relax [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Wed Sep 9 23:01:18 2026 -0700

    xfs: use correct jiffies comparison function in xchk_maybe_relax
    
    commit 984aab2d905a8557fafb27cd9e8713d6d12b3437 upstream.
    
    LOLLM points out that we're supposed to use time_after_eq, not a raw >=
    operation here, or else jiffies wraps can go unnoticed.  Fix this.
    
    Cc: [email protected] # v6.10
    Fixes: 271557de7cbfde ("xfs: reduce the rate of cond_resched calls inside scrub")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

xfs: use the correct reservations for rtrmap/refcount recovery [+ + +]
Author: Darrick J. Wong <[email protected]>
Date:   Thu Sep 10 22:54:11 2026 -0700

    xfs: use the correct reservations for rtrmap/refcount recovery
    
    commit 471e0b6e2ddac9e16b8dc2153d6e5575fefb8a2f upstream.
    
    LOLLM noticed that we might reserve the wrong number of blocks for
    recovering rtrmap and rtrefcount updates after a crash.  Fix that.
    
    Cc: [email protected] # v6.14
    Fixes: 5e0679d1c62f25 ("xfs: support recovering rmap intent items targetting realtime extents")
    Signed-off-by: Darrick J. Wong <[email protected]>
    Assisted-by: LOLLM # finding obvious bugs
    Reviewed-by: Christoph Hellwig <[email protected]>
    Signed-off-by: Carlos Maiolino <[email protected]>
    Signed-off-by: Greg Kroah-Hartman <[email protected]>

 
xsk: Use a 32-bit compare in xsk_map_gen_lookup [+ + +]
Author: Zhiling Zou <[email protected]>
Date:   Fri Sep 11 00:18:24 2026 +0800

    xsk: Use a 32-bit compare in xsk_map_gen_lookup
    
    [ Upstream commit 70504de0bb627848667207bec7ccfd647deb8814 ]
    
    xsk_map_gen_lookup() loads a u32 key and compares it with max_entries
    using BPF_JMP_IMM. BPF immediates are sign-extended to 64 bits, so a
    max_entries value of 0x80000000 or higher becomes a threshold larger
    than every zero-extended 32-bit key. An out-of-range index then skips
    the bounds check and the generated lookup reads past xsk_map[].
    
    Compare with BPF_JMP32_IMM so the check stays in 32-bit unsigned range.
    
    Fixes: e65650f291ee ("bpf: Implement map_gen_lookup() callback for XSKMAP")
    Reported-by: Vega <[email protected]>
    Signed-off-by: Zhiling Zou <[email protected]>
    Signed-off-by: Alexei Starovoitov <[email protected]>
    Reviewed-by: Emil Tsalapatis <[email protected]>
    Link: https://patch.msgid.link/7d2cb8e8dfaa9eb8fdff85156987a60960787dc3.1789056660.git.zhilinz@nebusec.ai
    Signed-off-by: Eduard Zingerman <[email protected]>
    Signed-off-by: Sasha Levin <[email protected]>