LLVM/project 99bab63 — llvm/lib/Target/AMDGPU GCNHazardRecognizer.cpp, llvm/test/CodeGen/AMDGPU wmma-coexecution-valu-hazards.mir

[AMDGPU] Fix missed WMMA C-operand co-exec hazard

The gfx1250 WMMA co-execution hazard check treats only A, B and the
SWMMAC index as registers the in-flight MMA still reads. C (src2 of a
non-SWMMAC WMMA) is missing, so a VALU scheduled into the MMA's shadow
can clobber C and the MMA consumes the new value.

This is latent while C is tied to vdst, since the existing D check then
covers it. It miscompiles where the tie does not hold: for
v_wmma_bf16f32_16x16x32_bf16, whose D is narrower than C, and for the
_threeaddr form of any WMMA.

Add src2 to the checked set for non-SWMMAC WMMAs.
DeltaFile
+171-2llvm/test/CodeGen/AMDGPU/wmma-coexecution-valu-hazards.mir
+5-0llvm/lib/Target/AMDGPU/GCNHazardRecognizer.cpp
+176-22 files

LLVM/project 20a018a — utils/bazel/llvm-project-overlay/llvm BUILD.bazel

[bazel] Use `includes` for per-target lib/Target dirs to fix -Wmicrosoft-include (#228282)

This is a cleaner implementation to
https://github.com/llvm/llvm-project/pull/227867 which provides a more
accurate-to-cmake include path instead of silencing the error.
DeltaFile
+14-11utils/bazel/llvm-project-overlay/llvm/BUILD.bazel
+14-111 files

LLVM/project 38d966f — compiler-rt/lib/sanitizer_common sanitizer_offload_hsa.h sanitizer_offload.h

[compiler-rt] Add more common HSA utilities for memory and symbolization

Summary:
Adds memory pool management for different kinds of allocation. Intended
to be used for the in-progress concurrency sanitizer.

RFC: discourse.llvm.org/t/rfc-a-thread-concurrency-sanitizer-for-gpus-in-compiler-rt/91113
DeltaFile
+46-3compiler-rt/lib/sanitizer_common/sanitizer_offload.cpp
+21-0compiler-rt/lib/sanitizer_common/sanitizer_offload_image.cpp
+15-0compiler-rt/lib/sanitizer_common/sanitizer_symbolizer_libcdep.cpp
+4-0compiler-rt/lib/sanitizer_common/sanitizer_symbolizer.h
+3-0compiler-rt/lib/sanitizer_common/sanitizer_offload.h
+2-0compiler-rt/lib/sanitizer_common/sanitizer_offload_hsa.h
+91-36 files

LLVM/project 36eb4f7 — compiler-rt/lib/csan csan_gpu.cpp csan.cpp, compiler-rt/test/csan volatile.cpp

Remove separate volatile entry points
DeltaFile
+49-0compiler-rt/test/csan/volatile.cpp
+22-0compiler-rt/test/csan/AMDGPU/volatile.hip
+0-4compiler-rt/lib/csan/csan_gpu.cpp
+0-4compiler-rt/lib/csan/csan.cpp
+71-84 files

LLVM/project 2a1a531 — compiler-rt/lib/csan csan.cpp, compiler-rt/test/csan access-sizes.cpp

Use unaligned loads when snapshotting watched values
DeltaFile
+22-0compiler-rt/test/csan/access-sizes.cpp
+3-3compiler-rt/lib/csan/csan.cpp
+25-32 files

LLVM/project 70b3ee3 — compiler-rt/cmake config-ix.cmake, compiler-rt/cmake/Modules AllSupportedArchDefs.cmake

comments
DeltaFile
+13-9compiler-rt/lib/csan/csan_gpu.cpp
+3-3compiler-rt/cmake/config-ix.cmake
+2-2compiler-rt/lib/csan/CMakeLists.txt
+1-1compiler-rt/test/csan/CMakeLists.txt
+1-1compiler-rt/lib/csan/offload/CMakeLists.txt
+1-0compiler-rt/cmake/Modules/AllSupportedArchDefs.cmake
+21-166 files

LLVM/project af23a10 — compiler-rt/lib/csan csan_watch.h csan_report.cpp, compiler-rt/lib/csan/offload csan_offload_report.cpp csan_offload_hsa_interceptors.cpp

[compiler-rt] Add 'csan' library for the concurrency sanitizer

Summary:
Adds the runtime for the concurrency sanitizer, both CPU and GPU.
Fundamentally, this works using the following pseudocode:

```c
static u64 watchpoints[N]; // Hash-indexed, zero is empty.

// Emitted before the access, so we never trip on our own write.
void check_access(volatile void *addr, u32 size, u32 type) {
    // Every access probes. A read conflicts only with a watched write, a
    // write conflicts with either.
    if (u64 *wp = find_watchpoint(addr, size, type))
        consume(wp, this_pc()); // Hand our location to the owner.

    if (!should_sample()) // Wave-uniform, 1-in-N chance.
        return;


    [17 lines not shown]
DeltaFile
+456-0compiler-rt/lib/csan/csan_gpu.cpp
+370-0compiler-rt/lib/csan/offload/csan_offload_hsa_interceptors.cpp
+334-0compiler-rt/lib/csan/csan.cpp
+279-0compiler-rt/lib/csan/csan_report.cpp
+184-0compiler-rt/lib/csan/offload/csan_offload_report.cpp
+163-0compiler-rt/lib/csan/csan_watch.h
+1,786-044 files not shown
+2,940-250 files

LLVM/project e933091 — compiler-rt/lib/csan CMakeLists.txt, compiler-rt/lib/csan/offload csan_offload_preinit.cpp CMakeLists.txt

[compiler-rt] Remove dlsym interceptor and support `-shared-libsan` for CSan

Summary:
Follow the UBSan offload runtime. Offload now resolves HSA through the
global scope, so the `dlsym` interceptor is no longer needed. The real
HSA entry points are still taken from the loaded HSA library rather than
`RTLD_NEXT`, since every DSO with a static runtime exports the same
wrappers and they would otherwise chain back into each other.

Build `libclang_rt.csan.so` with the offload objects folded in. The
exported HSA wrappers report failure when HSA is absent and warn when HSA
was loaded ahead of the runtime. The preinit hook moves to a separate
`csan_offload-preinit` archive for executables.
DeltaFile
+34-98compiler-rt/lib/csan/offload/csan_offload_hsa_interceptors.cpp
+84-9compiler-rt/lib/csan/CMakeLists.txt
+35-0compiler-rt/test/csan/hsa-load-order.cpp
+22-4compiler-rt/lib/csan/offload/CMakeLists.txt
+24-0compiler-rt/test/csan/hsa-missing.cpp
+22-0compiler-rt/lib/csan/offload/csan_offload_preinit.cpp
+221-1115 files not shown
+241-11111 files

LLVM/project 9d8964f — compiler-rt/lib/sanitizer_common sanitizer_offload_hsa.h sanitizer_offload.h

[compiler-rt] Add device memory pool and data symbolization to offload layer

Summary:
This adds a few helpers to the common sanitizer offload layer that the
GPU concurrency sanitizer will need. `Offload::GetMemoryPool` and
`Offload::Allocate` allocate device-local memory from an agent's
coarse-grained memory pool. `Offload::SymbolizeData` and
`Symbolizer::SymbolizeModuleData` symbolize global data addresses inside
device images, which are not mapped into the host process.
DeltaFile
+46-3compiler-rt/lib/sanitizer_common/sanitizer_offload.cpp
+21-0compiler-rt/lib/sanitizer_common/sanitizer_offload_image.cpp
+15-0compiler-rt/lib/sanitizer_common/sanitizer_symbolizer_libcdep.cpp
+4-0compiler-rt/lib/sanitizer_common/sanitizer_symbolizer.h
+3-0compiler-rt/lib/sanitizer_common/sanitizer_offload.h
+2-0compiler-rt/lib/sanitizer_common/sanitizer_offload_hsa.h
+91-36 files

LLVM/project d63124f — flang/include/flang/Evaluate tools.h, flang/lib/Evaluate tools.cpp

[flang][cuda] Look through associate names for managed and unified data (#228206)

IsCUDADeviceSymbol looks through an associate name: the name is device
data when its selector has device symbols. The managed and unified
predicates only handled object entities, so an associate name whose
selector is managed data was counted as device data that is not managed.

In host code, an element assignment such as
```
  associate(px => g%x, py => g%y)
    px%a(i,j) = r + px%s * real(py%n, 8)
  end associate
```

where `a` and `y` are managed, was then classified as a data transfer.
This
emitted a `cuf.data_transfer` from a scalar value, which the verifier
rejects.
IsCUDADataAttrSymbol now gives an associate name the attribute of the

    [2 lines not shown]
DeltaFile
+67-0flang/test/Lower/CUDA/cuda-associate-data-transfer.cuf
+56-5flang/lib/Evaluate/tools.cpp
+19-18flang/include/flang/Evaluate/tools.h
+142-233 files

NetBSD/pkgsrc 5ZeGbPC — doc TODO

   doc/TODO: update two

   - apache-2.4.69.

   + thrift-0.25.0 [CVE-2026-41608].
VersionDeltaFile
1.28068+2-3doc/TODO
+2-31 files

LLVM/project 5263866 — clang/lib/CIR/CodeGen CIRGenFunction.cpp, clang/test/CIR/CodeGenHLSL cxx-this-lvalue.hlsl

[CIR] Support CXXThisExpr in emitLValue (#227316)

Support `CXXThisExpr` in `CIRGenFunction::emitLValue` by wrapping
`loadCXXThisAddress()` into an LValue via `makeAddrLValue()`, matching
classic Clang codegen (`CGExpr.cpp:1865`).

This enables LValue contexts for `this`, such as member access
expressions via `this.field` in languages like HLSL where `this` is
reference-like.

Fixes #227193
DeltaFile
+61-0clang/test/CIR/CodeGenHLSL/cxx-this-lvalue.hlsl
+1-2clang/lib/CIR/CodeGen/CIRGenFunction.cpp
+62-22 files

NetBSD/pkgsrc HpOX2Sr — doc CHANGES-2026

   doc: Updated www/apache24 to 2.4.69
VersionDeltaFile
1.6580+2-1doc/CHANGES-2026
+2-11 files

NetBSD/pkgsrc EQG5dFk — www/apache24 Makefile distinfo

   www/apache24: update to 2.4.69

   Apache 2.4.69 (2026-10-01)

     *) Fix the tar icon in the documentation so that its background is
        transparent.  #70238.  [Jeffery To <jeffery.to gmail.com>]

     *) mod_ssl: Fix OpenSSL compatibility macros for X509_get0_notBefore,
        X509_get0_notAfter, and X509_get0_serialNumber with OpenSSL < 1.1.
        #70205.  [Craig Lorentzen <crlorent amazon.com>]

     *) mod_cgid, mod_ssl, mod_md: Various hardening fixes.  [Various authors]

     *) mod_md: MDServerStatus is now disabled by default.  [Joe Orton]

     *) mod_auth_digest: Fix compatibility with expression-based AuthName.
        #59039.  [Eric Covener]

     *) mod_auth_digest.c: Drop RFC 2069 support; rewrite shared memory

    [29 lines not shown]
VersionDeltaFile
1.41+10-1www/apache24/PLIST
1.73+4-4www/apache24/distinfo
1.149+2-3www/apache24/Makefile
+16-83 files

LLVM/project 79f4151 — flang/lib/Semantics check-cuda.h check-cuda.cpp, flang/test/Semantics/CUDA cuf-device-data-host-read.cuf

Revert "[flang][cuda] Diagnose host reads of device data" (#228286)

Reverts llvm/llvm-project#228271

auto-merging was on by mistake
DeltaFile
+0-133flang/lib/Semantics/check-cuda.cpp
+0-126flang/test/Semantics/CUDA/cuf-device-data-host-read.cuf
+0-11flang/lib/Semantics/check-cuda.h
+0-2703 files

LLVM/project b08cb38 — llvm/lib/CodeGen AtomicExpandPass.cpp, llvm/lib/Target/NVPTX NVPTXISelLowering.cpp

[AtomicExpand] Implement SUB → ADD(-x) and FSUB → FADD(-x) (#221425)

Implement ATOMIC_SUB → ATOMIC_ADD(-x) and ATOMIC_FSUB → ATOMIC_FADD(-x).
Many targets have atomic adds, but I don't think any have atomic subs.
This transform prevents cmpxchg expansion of atomic SUB. Enable these
transformations in NVPTX.

The integer variant exists in many different places currently. For
example, NVPTX currently expands it in DAG legalization.

begin AI generated

- SelectionDAG generic legalization:
[LegalizeDAG.cpp:3410](https://github.com/llvm/llvm-project/blob/4977a0c815c39fd90bb67c270174f2356ee9b5a7/llvm/lib/CodeGen/SelectionDAG/LegalizeDAG.cpp#L3410)
- GlobalISel generic legalization:
[LegalizerHelper.cpp:5117](https://github.com/llvm/llvm-project/blob/4977a0c815c39fd90bb67c270174f2356ee9b5a7/llvm/lib/CodeGen/GlobalISel/LegalizerHelper.cpp#L5117)
- GlobalISel outlined-atomic libcall path, mapping SUB to negated
`LDADD`:
[LegalizerHelper.cpp:919](https://github.com/llvm/llvm-project/blob/4977a0c815c39fd90bb67c270174f2356ee9b5a7/llvm/lib/CodeGen/GlobalISel/LegalizerHelper.cpp#L919)

    [42 lines not shown]
DeltaFile
+46-211llvm/test/CodeGen/NVPTX/atomicrmw-sm90.ll
+41-181llvm/test/CodeGen/NVPTX/atomicrmw-sm70.ll
+96-0llvm/test/Transforms/AtomicExpand/NVPTX/atomicrmw-sub.ll
+21-53llvm/test/CodeGen/NVPTX/atomicrmw-sm60.ll
+26-14llvm/lib/Target/NVPTX/NVPTXISelLowering.cpp
+26-0llvm/lib/CodeGen/AtomicExpandPass.cpp
+256-4591 files not shown
+260-4617 files

LLVM/project 6c1d106 — clang/lib/CIR/CodeGen CIRGenAsm.cpp CIRGenExpr.cpp, clang/lib/CIR/Dialect/IR CIRDialect.cpp

[CIR] Propagate the record address space to get_member (#226650)

Addresses: https://github.com/llvm/llvm-project/issues/226629

A member lives in its record's address space, but a few `get_member`
builders always produced a default-AS pointer. On SPIR-V that's private,
so a SYCL kernel was reading its captured pointer through a private
pointer. This patch takes the AS from the base and adds a verifier check
so we catch any stragglers.

Assisted-by: Claude / Opus 5.5
DeltaFile
+71-0clang/test/CIR/CodeGen/get-member-addrspace.cpp
+24-0clang/test/CIR/CodeGenHIP/inline-asm-multi-output-addrspace.hip
+14-0clang/test/CIR/IR/invalid-struct.cir
+6-3clang/lib/CIR/CodeGen/CIRGenExpr.cpp
+3-1clang/lib/CIR/CodeGen/CIRGenAsm.cpp
+3-0clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+121-43 files not shown
+125-79 files

LLVM/project b55367c — clang/lib/CIR/CodeGen CIRGenModule.h CIRGenDecl.cpp, clang/test/CIR/CodeGen amdgpu-array-addrspace.cpp

[CIR] Cast global addresses to their declared address space (#226649)

Opened to address a portion of
https://github.com/llvm/llvm-project/issues/226629


In CUDA, `__shared__ int sh` has type `int` but lives in AS 3. Classic
codegen casts the address to the declared type's AS where it's formed,
so users just see a generic pointer. We weren't doing that, so things
like `return &sh;` bitcast the slot instead, and NVPTX never got a
`cvta.shared`. This patch does the same cast in `getAddrOfGlobalVar` and
wherever static locals are fetched.

This also drops the comment claiming lowering would emit the cast for
us. That's only true for OpenCL, where the declared type already carries
the AS. LowerToLLVM never inserts casts on its own.

Assisted-by: Claude / Opus 5.5
DeltaFile
+95-0clang/test/CIR/CodeGenCUDA/global-addrspace-cast.cu
+26-11clang/test/CIR/CodeGen/amdgpu-array-addrspace.cpp
+18-3clang/lib/CIR/CodeGen/CIRGenModule.cpp
+3-9clang/lib/CIR/CodeGen/CIRGenDecl.cpp
+3-2clang/test/CIR/CodeGenCUDA/address-spaces.cu
+4-0clang/lib/CIR/CodeGen/CIRGenModule.h
+149-251 files not shown
+151-267 files

LLVM/project 560b66c — llvm/test/CodeGen/AMDGPU rewrite-vgpr-mfma-to-agpr-spill-multi-store-codegen.ll llvm.exp10.f64.ll

[AMDGPU] Retire dying uses before adding defs in GCNDownwardRPTracker

GCNDownwardRPTracker::advanceToNext() added an instruction's defs to the
live set while the uses that die at that instruction were still in it,
and only dropped them in the following advanceBeforeNext(). This could
lead to inflated register pressure reporting.

Apply the transfer function in the order the generic tracker uses in
RegPressureTracker::advance(): retire the lanes that die at the
instruction first, then add its defs. Early-clobber defs are the
exception. PHIs are skipped because nothing dies there. All of these are
now done by advanceToNext().

advanceBeforeNext() now only drops dead def lanes, the uses having been
retired already.

Two callers depend on the previous ordering and are updated:

- SIFormMemoryClauses::checkPressure() deliberately keeps dying uses live,

    [31 lines not shown]
DeltaFile
+9,499-9,281llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+3,117-3,106llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+594-646llvm/test/CodeGen/AMDGPU/load-global-i16.ll
+473-473llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.640bit.ll
+449-449llvm/test/CodeGen/AMDGPU/llvm.exp10.f64.ll
+400-424llvm/test/CodeGen/AMDGPU/rewrite-vgpr-mfma-to-agpr-spill-multi-store-codegen.ll
+14,532-14,37954 files not shown
+20,030-19,95760 files

LLVM/project edce25d — clang/lib/Driver/ToolChains CommonArgs.cpp, clang/test/Driver fsanitize.c fsanitize-concurrency-offload.c

[Clang] Support `-shared-libsan` for offload CSan

Summary:
Follow the UBSan handling. The shared runtime embeds the HSA
interceptors, so `csan_offload` is only linked with static runtimes and
executables pull in `csan_offload-preinit` to initialize early.
DeltaFile
+38-0clang/test/Driver/fsanitize-concurrency-offload.c
+9-4clang/lib/Driver/ToolChains/CommonArgs.cpp
+3-1clang/test/Driver/fsanitize.c
+50-53 files

LLVM/project 222e13f — clang/docs ConcurrencySanitizer.md, clang/lib/CodeGen CodeGenFunction.cpp

[Clang] Add support for the `-fsanitize=concurrency` runtime

Summary:
Add the frontend sanitizer kind, function attributes, pass pipeline
integration, predefined macro, driver handling, and documentation for
ConcurrencySanitizer.
DeltaFile
+207-0clang/docs/ConcurrencySanitizer.md
+73-0clang/test/CodeGen/sanitize-concurrency.c
+32-8clang/lib/Driver/ToolChains/CommonArgs.cpp
+15-0clang/test/Driver/fsanitize.c
+10-4clang/lib/Driver/SanitizerArgs.cpp
+10-3clang/lib/CodeGen/CodeGenFunction.cpp
+347-1510 files not shown
+375-1616 files

LLVM/project e8f3025 — clang/test/CodeGen sanitize-concurrency.c, clang/test/CodeGen/Inputs sanitize-concurrency-ignorelist.txt

ignore
DeltaFile
+29-0clang/test/CodeGen/sanitize-concurrency.c
+2-0clang/test/CodeGen/Inputs/sanitize-concurrency-ignorelist.txt
+31-02 files

LLVM/project c2c8c81 — clang/lib/Driver SanitizerArgs.cpp, clang/lib/Driver/ToolChains AMDGPU.cpp CommonArgs.cpp

Comments
DeltaFile
+48-0clang/test/Driver/fsanitize-concurrency-offload.c
+27-0clang/test/Driver/fsanitize.c
+16-7clang/lib/Driver/ToolChains/Clang.cpp
+11-4clang/lib/Driver/SanitizerArgs.cpp
+5-6clang/lib/Driver/ToolChains/CommonArgs.cpp
+3-2clang/lib/Driver/ToolChains/AMDGPU.cpp
+110-195 files not shown
+112-2111 files

LLVM/project f390ab1 — llvm/lib/Transforms/Instrumentation ConcurrencySanitizer.cpp

enum class and return Res
DeltaFile
+20-14llvm/lib/Transforms/Instrumentation/ConcurrencySanitizer.cpp
+20-141 files

LLVM/project 69c89de — llvm/test/CodeGen/AMDGPU slp-int-to-fp.ll

update slp-int-to-fp.ll
DeltaFile
+324-391llvm/test/CodeGen/AMDGPU/slp-int-to-fp.ll
+324-3911 files

LLVM/project 1458734 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU narrow-int-to-bfloat.ll narrow-int-to-fp.ll

[AMDGPU] Do not price the extension of a loaded i40, i48 or i56 in int to fp casts

A widened constant or invariant load still pays it.
DeltaFile
+100-100llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+24-24llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+18-6llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+142-1303 files

LLVM/project dfe42ed — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU narrow-int-to-bfloat.ll narrow-int-to-fp.ll

few fixes
DeltaFile
+90-90llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+24-24llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+10-1llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+124-1153 files

LLVM/project 8f5f5ea — llvm/test/CodeGen/AMDGPU slp-int-to-fp.ll

update slp-int-to-fp.ll
DeltaFile
+128-170llvm/test/CodeGen/AMDGPU/slp-int-to-fp.ll
+128-1701 files

LLVM/project 1408c21 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU narrow-int-to-bfloat.ll narrow-int-to-fp.ll

[AMDGPU] Price scalar integer to fp casts by source width and sign

Scalar sources between a byte and 31 bits fell to the default cost of one
while the matching vector lanes were already priced, which skewed the
difference SLP weighs a bundle against. Such a source is extended before
the conversion, and what the extension takes depends on the width, on the
sign and on whether the subtarget has SDWA and 16 bit instructions.
Sources narrower than a byte are left alone, because their vector form is
not priced either.
DeltaFile
+388-388llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+72-72llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+29-9llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+489-4693 files

LLVM/project 1151a47 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp

format
DeltaFile
+1-1llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+1-11 files