LLVM/project 134716f — clang/lib/CIR/CodeGen CIRGenExpr.cpp, clang/test/CIR/CodeGen array-to-pointer-decay.cpp

[CIR] Fix array-to-pointer decay (#230276)

Fixes #229654
DeltaFile
+30-0clang/test/CIR/CodeGen/array-to-pointer-decay.cpp
+4-9clang/lib/CIR/CodeGen/CIRGenExpr.cpp
+34-92 files

LLVM/project c0a1a2c — llvm/lib/Target/AMDGPU AMDGPULowerIntrinsics.cpp, llvm/test/CodeGen/AMDGPU llvm.amdgcn.wmma.f8f6f4.gfx1250.ll

[AMDGPU] Promote wmma_f32_16x16x128_f8f6f4 to scaled version on gfx1250-strict (#230253)

Added scale factor is 0 which does not change the result. "0" is a
special value that maps to the exponent "1.0" when passed as a constant.
DeltaFile
+37-0llvm/lib/Target/AMDGPU/AMDGPULowerIntrinsics.cpp
+35-0llvm/test/CodeGen/AMDGPU/llvm.amdgcn.wmma.f8f6f4.gfx1250.ll
+72-02 files

LLVM/project 09b471a — clang/lib/CIR/Dialect/Transforms CallConvLoweringPass.cpp, clang/test/CIR/CodeGen call-conv-lowering-x86_64-vec3.c

[CIR] Lower non-power-of-two vectors in CallConvLowering (#230087)

Summary:

- The x86_64 classifier now sizes a vector at its ABI size (#227889), so
the bridge no longer needs to reject vectors whose width is not a power
of two.
- The SSEUP walk rounds vector widths the same way, so a union holding
such a vector is still refused at sizes with no coerce type.
- This issue came up while enabling `CallConvLowering` for AMDGPU in
#220197: the gfx950/gfx1250 `transpose-load` builtins return
three-element vectors, and
`CIR/CodeGenHIP/builtins-amdgcn-gfx950-read-tr.hip` and
`builtins-amdgcn-gfx1250-load-tr.hip` fail on the gate.

Related to issue: #220471 

Assisted by : claude opus 5.5
DeltaFile
+86-0clang/test/CIR/CodeGen/call-conv-lowering-x86_64-vec3.c
+11-10clang/test/CIR/Transforms/abi-lowering/x86_64-aggregate-nyi.cir
+2-6clang/lib/CIR/Dialect/Transforms/CallConvLoweringPass.cpp
+99-163 files

LLVM/project f071a20 — clang/include/clang/AST Expr.h, clang/include/clang/Options Options.td

[Clang] Mark scoped_atomics with !noalias.addrspace(private)

The HIP specification marks atomics on thread private memory as UB.
Scoped atomics used within a HIP context are also considered UB,
unless explicitley specified via a command line argument.
These are now annotated with !noalias.addrspace(5) for amdgpus,
to avoid an expensive runtime check.
DeltaFile
+76-75clang/test/CodeGen/scoped-atomic-ops.c
+22-20clang/test/CodeGenCUDA/atomic-options.hip
+27-0clang/test/CodeGenCUDA/private-atomics-undefined.hip
+7-1clang/lib/CodeGen/Targets/AMDGPU.cpp
+3-4clang/include/clang/AST/Expr.h
+6-0clang/include/clang/Options/Options.td
+141-1002 files not shown
+143-1018 files

LLVM/project 7dd738d — llvm/include/llvm/CodeGen RegisterPressure.h, llvm/lib/CodeGen RegisterPressure.cpp

[NFC][RegisterPressure] Add RegisterOperands::restoreLivenessFlags helper (#229952)

Factor the
clear-stale-read-undef-flags-then-recompute-from-LiveIntervals pattern,
currently duplicated in `GCNIterativeScheduler::restoreLivenessFlags`
and `GCNSchedStrategy::modifyRegionSchedule`, into a shared
`RegisterOperands::restoreLivenessFlags` helper, and convert both AMDGPU
call sites to it with no behavior change. The helper also accepts an
optional register filter for targeted recomputation.

Requested as a prerequisite refactor in review of #227897, which will
stack on top of this and switch its new helper to the shared one.
DeltaFile
+30-0llvm/lib/CodeGen/RegisterPressure.cpp
+3-14llvm/lib/Target/AMDGPU/GCNIterativeScheduler.cpp
+3-11llvm/lib/Target/AMDGPU/GCNSchedStrategy.cpp
+14-0llvm/include/llvm/CodeGen/RegisterPressure.h
+0-2llvm/lib/Target/AMDGPU/GCNIterativeScheduler.h
+50-275 files

LLVM/project 76f1093 — llvm/lib/Support UnicodeNameToCodepointGenerated.cpp, llvm/test/CodeGen/AMDGPU amdgcn.bitcast.832bit.ll amdgcn.bitcast.896bit.ll

Merge branch 'users/mssefat/anti-hints-pr4-amdgpu-pre-ra' into users/mssefat/anti-hints-pr5-amdgpu-pre-ra-coexec-anti-hint
DeltaFile
+75,746-95,162llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+23,347-23,371llvm/lib/Support/UnicodeNameToCodepointGenerated.cpp
+20,168-25,874llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+16,417-20,023llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.960bit.ll
+15,309-18,740llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.896bit.ll
+14,068-17,405llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.832bit.ll
+165,055-200,57510,020 files not shown
+726,250-533,75910,026 files

LLVM/project be2daae — llvm/lib/Target/RISCV RISCVInstrInfoXqccmp.td RISCVInstrInfoZc.td, llvm/lib/Target/RISCV/Disassembler RISCVDisassembler.cpp

[RISCV] Reject reserved (qc.)cm.mvsa01/(qc.)cm.mva01s encodings in the disassembler (#229091)

The disassembler accepted every value of the two 3-bit sreg fields of
`(qc.)cm.mvsa01` and `(qc.)cm.mva01s`, but the Xqccmp/Zcmp specification
reserves some of them:

- (qc.)cm.mvsa01 requires r1s' != r2s'. Encodings with equal fields are
reserved:
- Zcmp:
https://github.com/riscv/riscv-isa-manual/blob/1e0debba4a909dba7cce7a0e6879fd79b3b6e2b9/src/unpriv/zcmp.adoc#L1160-L1163
- Xqccmp:
https://github.com/qualcomm/riscv-unified-db/blob/Xqccmp_extension-0.3.0/arch_overlay/qc_iu/inst/Xqccmp/qc.cm.mvsa01.yaml#L8

- On RVE only s0 and s1 exist, so an sreg field greater than 1 is
reserved for both instructions:
Zcmp:
https://github.com/riscv/riscv-isa-manual/blob/1e0debba4a909dba7cce7a0e6879fd79b3b6e2b9/src/unpriv/zcmp.adoc#L1197-L1199
Zcmp:
https://github.com/riscv/riscv-isa-manual/blob/1e0debba4a909dba7cce7a0e6879fd79b3b6e2b9/src/unpriv/zcmp.adoc#L1265-L1267
DeltaFile
+63-0llvm/test/MC/Disassembler/RISCV/zcmp-mv.txt
+63-0llvm/test/MC/Disassembler/RISCV/xqccmp-mv.txt
+16-2llvm/lib/Target/RISCV/Disassembler/RISCVDisassembler.cpp
+7-1llvm/lib/Target/RISCV/RISCVInstrInfoZc.td
+2-1llvm/lib/Target/RISCV/RISCVInstrInfoXqccmp.td
+151-45 files

LLVM/project e1df86a — clang/test/CIR/CodeGen call-conv-lowering-x86_64-vec3.c

test update
DeltaFile
+4-4clang/test/CIR/CodeGen/call-conv-lowering-x86_64-vec3.c
+4-41 files

LLVM/project 33840cf — clang/include/clang/CIR/Dialect/IR CIROps.td, clang/lib/CIR/Dialect/IR CIRDialect.cpp

[CIR] Implement BranchOpInterface for switch.flat (#230340)

cir.switch.flat is the CFG-form terminator produced by CIR flattening,
but unlike cir.br and cir.brcond, it doesn't implmenet BranchOpIterface.
As a result, generic MLIR control-flow infrastructure could not tell
which operands are forwarded to which successor's block arguments,. For
example, SCCP pass. When the switch condition is a know constant, SCCP
can know that which sucessor block is actually taken. Before, SCCP
didn't understand switch.flat. It treated every sucessor as reachable,
which forbid it to fold anything.

Implmenet the interface the same way llvm.switch do now.

Assisted-by: Claude # Tests
DeltaFile
+70-0clang/test/CIR/Transforms/switch-flat-sccp.cir
+20-0clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+4-4clang/test/CIR/IR/switch-flat.cir
+1-0clang/include/clang/CIR/Dialect/IR/CIROps.td
+95-44 files

LLVM/project eb95172 — flang/lib/Optimizer/Transforms/CUDA CUFDeviceGlobal.cpp, flang/test/Fir/CUDA cuda-device-global.f90

[flang][cuda] Copy type descriptors for array descriptors in device code (#230305)

The CUFDeviceGlobal pass copies the type descriptors of derived types
used
in device procedures into the GPU module. For fir.embox, it only looked
at
the memref type after stripping the reference, so a scalar derived type
was
handled but an array of derived type (e.g. passing the section `a(1:1)`
of
an assumed-size array to an assumed-shape dummy) was not. The type
descriptor was then missing from the GPU module and FIR-to-LLVM codegen
failed with "runtime derived type info descriptor was not generated".

Use the element type of the resulting box to find the derived type,
which
covers both scalars and arrays. Handle fir.rebox the same way, since
reboxing a non-polymorphic derived type also needs the type descriptor
in
codegen.
DeltaFile
+34-0flang/test/Fir/CUDA/cuda-device-global.f90
+9-5flang/lib/Optimizer/Transforms/CUDA/CUFDeviceGlobal.cpp
+43-52 files

LLVM/project 487f7f9 — clang/lib/CodeGen BackendUtil.cpp, llvm/include/llvm/Passes RunCodeGen.h

[Passes] Remove unused PrintPipelinePasses parameter from runCodeGenPipeline. NFC (#230349)
DeltaFile
+9-9llvm/lib/Passes/RunCodeGen.cpp
+2-3clang/lib/CodeGen/BackendUtil.cpp
+1-2llvm/include/llvm/Passes/RunCodeGen.h
+12-143 files

LLVM/project d7a6e95 — clang-tools-extra/clangd/unittests FindTargetTests.cpp, clang/docs ReleaseNotes.md

[Clang] Still track rewrite operator== from the spaceship operator (#230062)

This reverts the behavior since d4cf20ca37160cb062a9db773d0e6255d6bbc31a

We still need to track the instantiated-from decl for rewritten
operators, in order to suppress spurious warnings that suggest the
spaceship was not used.

Also this fixes regressions introduced from that commit.

Fixes https://github.com/llvm/llvm-project/issues/125233
Fixes https://github.com/llvm/llvm-project/issues/104720
DeltaFile
+52-5clang/test/CXX/class/class.compare/class.compare.default/p4.cpp
+11-13clang/lib/Sema/SemaTemplateInstantiateDecl.cpp
+4-1clang/lib/Sema/SemaConcept.cpp
+4-1clang-tools-extra/clangd/unittests/FindTargetTests.cpp
+4-0clang/lib/Sema/SemaOverload.cpp
+4-0clang/docs/ReleaseNotes.md
+79-201 files not shown
+81-227 files

LLVM/project 6be987a — llvm/lib/Support UnicodeNameToCodepointGenerated.cpp, llvm/test/CodeGen/AMDGPU amdgcn.bitcast.832bit.ll amdgcn.bitcast.896bit.ll

Merge remote-tracking branch 'upstream/main' into users/mssefat/anti-hints-pr4-amdgpu-pre-ra
DeltaFile
+75,746-95,162llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+23,347-23,371llvm/lib/Support/UnicodeNameToCodepointGenerated.cpp
+20,168-25,874llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+16,417-20,023llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.960bit.ll
+15,309-18,740llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.896bit.ll
+14,068-17,405llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.832bit.ll
+165,055-200,57510,020 files not shown
+726,264-533,76510,026 files

LLVM/project c4bbc6e — llvm/lib/Target/RISCV/GISel RISCVLegalizerInfo.cpp, llvm/test/CodeGen/RISCV/GlobalISel legalizer-info-validation.mir half-convert.ll

[RISCV][GlobalISel] Promote f16 to f32 for G_FPTOSI/G_FPTOUI without Zfh (#230025)

Use the fcvt.s.h conversion with Zfhmin and the __extendhfsf2 libcall
without it, then convert from f32, matching SelectionDAG. This also
fixes failures to legalize fptoui/fptosi from half to i32 without Zfh.
DeltaFile
+106-22llvm/test/CodeGen/RISCV/GlobalISel/half-convert.ll
+11-2llvm/lib/Target/RISCV/GISel/RISCVLegalizerInfo.cpp
+4-4llvm/test/CodeGen/RISCV/GlobalISel/legalizer-info-validation.mir
+121-283 files

LLVM/project 5b0ebe6 — clang/include/clang/CIR/Dialect/IR CIROps.td, clang/lib/CIR/Dialect/IR CIRDialect.cpp

[CIR] Implement BranchOpInterface for switch.flat

Switch.flat should minic llvm.switch as they are both CFG
formed. Implement the same mechanism as what we have in llvm.switch
and add ArrayRef<Attribute> override when we actually know the operands.

Assisted-by: Claude # Tests
DeltaFile
+70-0clang/test/CIR/Transforms/switch-flat-sccp.cir
+20-0clang/lib/CIR/Dialect/IR/CIRDialect.cpp
+4-4clang/test/CIR/IR/switch-flat.cir
+1-0clang/include/clang/CIR/Dialect/IR/CIROps.td
+95-44 files

LLVM/project 585e300 — clang/lib/CodeGen CGExpr.cpp CGExprAgg.cpp, clang/test/CodeGenObjC strong-in-c-struct.m

[Clang][CodeGen] Fix clang codegen eh cleanup (NFC) (#226743)

NFC for the following reasons:

- `DK_none` and `DK_cxx_destructor` can never reach
`pushLifetimeExtendedDestroy` at either of the two call sites. `DK_none`
is filtered out by the `if (...)` check. `DK_cxx_destructor ` only
occurs for a C++ class, but the code runs only if
`!getLangOpts().CPlusPlus`.

- For the other three (`DK_objc_strong_lifetime`,
`DK_objc_weak_lifetime`, and `DK_nontrivial_c_struct`), the boolean flag
controls whether the remaining not-yet-destroyed elements of an array
get destroyed in `emitArrayDestroy` when the destroyer call itself
throws when destroying an array element. But the flag is irrelevant here
because the destroyer functions for the three never throw: they are all
called via `EmitNounwindRuntimeCall`, which produces a plain `call`
rather than an `invoke`.
DeltaFile
+35-0clang/test/CodeGenObjC/strong-in-c-struct.m
+2-3clang/lib/CodeGen/CGExprAgg.cpp
+1-3clang/lib/CodeGen/CGExpr.cpp
+38-63 files

LLVM/project 19d3b2d — mlir/lib/Target/LLVMIR DebugTranslation.cpp

[mlir][NFC] Use DefaultUnreachable in TypeSwitch (#230333)

Simplify unreachable case in `TypeSwitch` to use `DefaultUnreachable`.
Similar to #162010.
DeltaFile
+2-3mlir/lib/Target/LLVMIR/DebugTranslation.cpp
+2-31 files

LLVM/project 3be74a9 — compiler-rt/lib/asan asan_rtl.cpp

Revert PrintAddressSpaceLayout changes, only print kGaplessShadow flag
DeltaFile
+36-41compiler-rt/lib/asan/asan_rtl.cpp
+36-411 files

LLVM/project 8b775e2 — compiler-rt/lib/asan asan_mapping.h asan_shadow_setup.cpp

[ASan][Darwin] Support gapless shadow layout for iOS 27.0

When the shadow can be placed entirely above app memory (as on the
new iOS 27.0 embedded VM layout, where debug memory pushes shadow
past kHighMemEnd), there is no need to split shadow into low/high
halves with a middle gap.

- Add kGaplessShadow (Apple-only) to detect this configuration.
- Teach InitializeShadowMemory to reserve one contiguous shadow
  region and protect only the shadow-of-shadow when kGaplessShadow
  is true, with CHECKs asserting the mapping preconditions.
- Update PrintAddressSpaceLayout to print the single-region layout.

rdar://167657399
DeltaFile
+41-35compiler-rt/lib/asan/asan_rtl.cpp
+39-1compiler-rt/lib/asan/asan_shadow_setup.cpp
+9-0compiler-rt/lib/asan/asan_mapping.h
+89-363 files

LLVM/project de95e26 — compiler-rt/lib/sanitizer_common sanitizer_common.h sanitizer_platform.h, compiler-rt/test/asan/TestCases/Darwin sandbox-vm-region-recurse.cpp

[sanitizer_common][Darwin] Add debug memory region support for iOS 27.0

iOS 27.0 bumps the address space from 36 to 39 bits on some devices,
and reserves some address space for sanitizers.

- Add SANITIZER_IOSDEVICE and SANITIZER_EMBEDDED_VM_LAYOUT macros;
  bump Darwin iOS/ARM64 SANITIZER_MMAP_RANGE_SIZE from 36 to 39 bits.
- Add ActivateDebugMemory / DebugMemoryActive and tag mmap allocations
  above DARWIN_DEBUG_MEMORY_START with VM_MEMORY_DEBUG when using the
  debug range; verify returned addresses fall within the expected range.
- Replace GetAppReservedRanges with GetAppRanges, populated from
  sysctls on supported devices.
- Extend FindAvailableMemoryRange with a use_debug_vm parameter and
  route MapDynamicShadow through it when debug memory is active.
- Centralize mach_vm_region_recurse calls through
  internal_mach_vm_region_recurse, which fatals on KERN_DENIED.

rdar://167657399
DeltaFile
+333-68compiler-rt/lib/sanitizer_common/sanitizer_mac.cpp
+15-15compiler-rt/lib/sanitizer_common/sanitizer_win.cpp
+19-1compiler-rt/lib/sanitizer_common/sanitizer_mac.h
+6-2compiler-rt/lib/sanitizer_common/sanitizer_platform.h
+0-4compiler-rt/lib/sanitizer_common/sanitizer_common.h
+1-1compiler-rt/test/asan/TestCases/Darwin/sandbox-vm-region-recurse.cpp
+374-916 files

LLVM/project 9429767 — bolt/lib/Rewrite RewriteInstance.cpp, bolt/test/AArch64 split-func-mapping-symbol.s

[BOLT][AArch64] Don't treat entry mapping symbols as function symbols (#230244)

ARM code mapping symbol ($x) at a function entry is treated as a
function alias. When rewriting, it's updated to match the function:
* gets function output size - while mapping syms should be size 0,
* fragment/ICF symbols are derived from it (e.g. $x.cold.0), without
  corresponding parent symbol.

Move mapping symbols together with the function but keep 0-sized and
don't add extra symbols for them.

This fixes issues of reading BOLTed binary with such symbols where
perf2bolt/heatmap report:
* "parent function not found for $x.cold.0",
* "owning FILE symbol not found for symbol $x.cold.0".

Test Plan: added split-func-mapping-symbol.s

Assisted-by: Claude Opus 5.5
DeltaFile
+71-0bolt/test/AArch64/split-func-mapping-symbol.s
+10-2bolt/lib/Rewrite/RewriteInstance.cpp
+81-22 files

LLVM/project 98ecca7 — clang/test/CIR/CodeGenSYCL sycl-module-id.cpp

[CIR][SYCL][NFC] Add sycl-module-id test (#230324)

Port clang/test/CodeGenSYCL/sycl-module-id.cpp to ClangIR.

FWIW I will port the SYCL tests from OG to CIR. That will likely make it
easier to track what is missing.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
DeltaFile
+35-0clang/test/CIR/CodeGenSYCL/sycl-module-id.cpp
+35-01 files

LLVM/project 561fa0e — clang/test/CIR/CodeGenSYCL functionptr-addrspace.cpp

[CIR][SYCL][NFC] Add functionptr-addrspace test (#230327)

Port clang/test/CodeGenSYCL/functionptr-addrspace.cpp to ClangIR.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
DeltaFile
+33-0clang/test/CIR/CodeGenSYCL/functionptr-addrspace.cpp
+33-01 files

LLVM/project 0ffb828 — bolt/lib/Rewrite RewriteInstance.cpp, bolt/test bolt-reserved.test

[BOLT] Ignore __bolt_reserved syms in perf2bolt/heatmap (#230300)

Aggregation and heatmap don't rewrite the binary, so skipping setting
reserved space is safe. This unblocks aggregation from BOLTed binary
that used and updated these markers.

Test Plan: updated bolt-reserved.test
DeltaFile
+13-0bolt/test/bolt-reserved.test
+4-0bolt/lib/Rewrite/RewriteInstance.cpp
+17-02 files

LLVM/project a26bff1 — llvm/lib/Target/WebAssembly WebAssemblyFastISel.cpp, llvm/test/CodeGen/WebAssembly fast-isel-simd128.ll

[WebAssembly] Fall back to SelectionDAG for vector compares in FastISel (#227368)

FastISel's selectICmp and selectFCmp each classify compare operands
with a single scalar test: anything but i64 is treated as i32, and
anything but f64 as f32, vectors included. A vector compare therefore
emits a scalar i32.eq/i32.ne (or f32.eq) that consumes the two v128
registers holding the operands, and the resulting module fails
validation ("type mismatch: expected i32, found v128").

Return false for vector operands in both selectors so the SelectionDAG
lowers them to SIMD compares, and extend the existing
fast-isel-simd128.ll test to cover vector compares on both wasm32 and
wasm64 (UTC-generated checks, with the vectors passed in as arguments).

The vector compare only reaches this path when the rest of the block
does not bail out, which is why it went unnoticed: the common shapes
(a vector icmp feeding extractelement, a bitcast, an intrinsic, or a
branch condition) all miss somewhere in FastISel and the SelectionDAG
revisits the block, replacing the bad compare. A vector compare

    [4 lines not shown]
DeltaFile
+65-0llvm/test/CodeGen/WebAssembly/fast-isel-simd128.ll
+14-0llvm/lib/Target/WebAssembly/WebAssemblyFastISel.cpp
+79-02 files

LLVM/project 825b7fa — libc/include/llvm-libc-macros/linux sys-inotify-macros.h, libc/include/sys inotify.yaml

[libc] add sys/inotify (#230314)

The sys/inotify header has linux functions for watching a directory.
This PR adds those functions, and their relevant macros and type. Tests
are fairly simple since the functions are all syscall wrappers.

Assisted-by: Automated tooling, human reviewed
DeltaFile
+82-0libc/include/sys/inotify.yaml
+73-0libc/test/src/sys/inotify/linux/CMakeLists.txt
+56-0libc/src/sys/inotify/linux/CMakeLists.txt
+54-0libc/include/llvm-libc-macros/linux/sys-inotify-macros.h
+50-0libc/test/src/sys/inotify/linux/inotify_init_test.cpp
+49-0libc/src/__support/OSUtil/linux/syscall_wrappers/CMakeLists.txt
+364-035 files not shown
+1,028-041 files

LLVM/project 8b7a4fb — llvm/utils/gn/secondary/clang/lib/CodeGenUtils BUILD.gn

[gn build] Port a789eb1d665b (#230326)
DeltaFile
+0-1llvm/utils/gn/secondary/clang/lib/CodeGenUtils/BUILD.gn
+0-11 files

LLVM/project b08e005 — clang/lib/CIR/CodeGen CIRGenCall.cpp, clang/test/CIR/CodeGen spir-call-calling-conv.c spir-call-calling-conv.cpp

[CIR] Emit calling convention on call sites (#230123)

Adds calling convention on the call sites. 

Note: I opted to execlude support for runtime calling convention
(classic getRuntimeCC()) on calls to runtime functions created in CIR
passes (EH, __cxa_atexit, __cxa_guard_*, dynamic_cast, global init).
These keep the default C calling convention and are marked with
TODO(cir) and `MissingFeatures::opFuncCallingConv()` for a follow-up.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
DeltaFile
+65-0clang/test/CIR/CodeGen/spir-call-calling-conv.cpp
+41-0clang/test/CIR/CodeGen/spir-call-calling-conv.c
+40-0clang/test/CIR/Lowering/call-calling-conv.cir
+39-0clang/test/CIR/IR/call-calling-conv.cir
+38-0clang/test/CIR/CodeGenSYCL/call-calling-conv.cpp
+21-16clang/lib/CIR/CodeGen/CIRGenCall.cpp
+244-1621 files not shown
+476-6627 files

LLVM/project f93f4e3 — llvm/lib/CodeGen/AsmPrinter EHStreamer.cpp, llvm/test/CodeGen/X86 eh-call-site-info-for-inline-asm.ll

[X86] Emit LSDA call site info for inline asm calls marked `unwind` (#218276)

LLVM currently doesn't consider inline asm calls with the `unwind`
keyword when generating the list of LSDA call sites of a function.
For example, given the following IR

```llvm
declare i32 @rust_eh_personality(i32, i32, i64, ptr, ptr)
declare void @bar()

define void @example() personality ptr @rust_eh_personality {
entry:
  call void asm sideeffect alignstack inteldialect unwind "call foo", ""()

  invoke void @bar()
          to label %cont unwind label %lpad

cont:
  ret void

    [98 lines not shown]
DeltaFile
+49-0llvm/test/CodeGen/X86/eh-call-site-info-for-inline-asm.ll
+6-0llvm/lib/CodeGen/AsmPrinter/EHStreamer.cpp
+55-02 files

LLVM/project 89157ec — bolt/include/bolt/Core BinaryContext.h, bolt/lib/Core BinaryEmitter.cpp BinaryContext.cpp

[BOLT] Keep ambiguous references next to function boundaries valid

A reference into code without a relocation that names its target, e.g. a
RIP-relative LEA whose relocation was against a section symbol, cannot be told
apart from "Next - Delta" and "Prev + Offset" when it lands right before or
after a function start. HHVM built with LLVM 23 has such a reference to
"sqlite3RCStrUnref - 1", which lands in padding and makes BOLT fail with
-strict. Without padding, BOLT silently kept such references relative to the
preceding function. See bolt/test/X86/unanchored-code-reference.s.

BOLT now collects these references from code and, when they are within
--boundary-ref-distance bytes (default 2) of a function start, keeps the
functions around them in place in lite mode. When all functions are processed,
it emits those functions unoptimized and back-to-back with the original bytes
between them, and checks after linking that they kept their size and distance.

Absolute references right before a function start are now relative to that
function. This replaces the workaround for "fptr - 1" in de-virtualized member
function pointer calls, which dropped the relocation and left the input
address in the code.
DeltaFile
+253-0bolt/test/X86/unanchored-code-reference.s
+224-0bolt/test/X86/unanchored-code-reference-fuse.s
+221-0bolt/lib/Core/BinaryContext.cpp
+57-21bolt/lib/Rewrite/RewriteInstance.cpp
+49-0bolt/include/bolt/Core/BinaryContext.h
+43-4bolt/lib/Core/BinaryEmitter.cpp
+847-253 files not shown
+883-259 files