Address review comments
- Rename the ACE tests and directories to acev1 for consistency
- Build acev1 on top of avx10.1 rather than avx10.2
- Fix test encodings and the BSR0 implicit defs
- Make the scale register type consistent with the spec
- Add memory operand tests for bsrmov
- Confirm the ACEV1 ISA implication chain
- Take the outer product ZMM sources by value instead of by register ID,
so they are allocated conventionally, and mark the 512-bit ACE builtins
with RequiredVectorWidth<512>
- Use pseudos for SPILL and RELOAD. Spills use scratch zmm. Don't use
non-ace ISA.
- Fix missing vector r/m in MRMSrcReg4VOp3.
[DirectX] Refactor DXILDebugInfo into a pass (#220989)
resolves #220323
DXILDebugInfo was structured as a namespace
with its pass chained to run as part of either
the DXILBitcodeWriter or DXILPrettyPrinter passes.
This is wrong and led to us failing expensive checks because
DXILDebugInfo changes IR but the writers are not expected to do so.
The fix was two fold, move the DXILDebugInfo into a pass. Second,
Refactor all the module debug info reader\writer code out of
DXILDebugInfo and into DXILDebugInfoMap which is used by both the
PrettyPrinter and the BitCodeWriter and move that code into the
DXILWriter compilation module since it is only used by the writers.
[Scalarizer][DirectX] Teach the Scalarizer to handle integer bitcasts between scalars and vectors. (#221033)
Fixes #219327
For vector-to-scalar bitcasts, combine the scalarized vector elements
using
extensions, shifts, and ORs. This avoids reconstructing an illegal
vector before
bitcasting it to a scalar. In particular, this fixes the `<2 x i16> ->
i32`
bitcast generated while lowering 16-bit vector `countbits`.
Also move the existing `i64 -> <2 x i32>` legalization from the DirectX
legalizer into the generic Scalarizer. The Scalarizer now splits these
values
using shifts and truncations while respecting target endianness.
With both bitcast directions handled by the Scalarizer, remove
`legalizeGetHighLowi64Bytes` from the DirectX legalizer and move
[5 lines not shown]
AMDGPU: Warn if trying to codegen without a subarch (#220245)
Warn if using the legacy amdgcn name, or amdgpu without a
specified subarch. This is to push all the non-clang frontends
to update to the new system, but this should turn into an error
in the next release.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
[CIR][NFC] Use declarative llvmOp lowering for coroutine intrinsics (#221343)
This PR switches the coroutine intrinsic ops (`cir.coro.intrinsic.id`,
`.alloc`, `.begin`, `.end`, `.free`, `.size`) to use the declarative
`llvmOp` field instead of hand-written lowering patterns.
[Analysis] Remove dead WriteDOTGraphToFile (NFC) (#221413)
The last callers were migrated to DOTGraphTraitsPrinter on May 16, 2022
in commit 7dce9eb6e507d48d0b79bfb408592936d378cc28.
Assisted-by: Antigravity
[CodeGen] Use RegisterClassInfo for remaining allocation-order users (#216510)
Greedy already uses RegisterClassInfo for the function-specific
allocation order (reserved registers filtered, CSRs deferred). Remaining
users still walked TRI's raw order or rebuilt RCI themselves.
Uses RCI.getOrder() everywhere that list is needed (PBQP, Hexagon
spill-slot search*, ARM load/store opt, AMDGPU AGPR-copy MFMA rewrite).
Takes the shared analysis when the pass can preserve it; otherwise keep
a
local RCI after freezeReservedRegs().
Hexagon findPhysReg no longer considers reserved registers.
**Note:** Initially had register scavenger use RCI.getOrder() for
function specific scratch reg allocation, but due to test failures and
the extra loops needed as fallbacks to fix this, this change ends up
likely being outside the scope of this PR so I removed it.
**Another note:** \*Hexagon changes removed from this PR, split into
#216752
ARM: Form fused VFMA/VFMS from the contract flag (#221340)
Select the fused VFMA/VFMS/VFNMA/VFNMS from the per-node contract
fast-math flag instead of the global AllowFPOpFusion == Fast. This is one of the
few remaining consumers of the TargetOption field.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[CIR][OpenCL] Lower OpenCL language version metadata to LLVM dialect
Propagate CIR OpenCL language version module attributes as LLVM dialect named metadata before LLVM IR translation.
Assisted-by: Codex / GPT-5.6 Sol
fix: Supply the HIP SPIR-V version only for metadata emission
Fix the assertion exposed by PR #214246 under the version invariant from PR #219687. Supply OpenCL 2.0 in classic CodeGen and CIRGen without changing HIP language options or enabling OpenCL-only Sema restrictions.
Assisted-by: Codex / GPT-6
[CIR][OpenCL] Emit OpenCL language version metadata in CIR
Emit OpenCL and C++ for OpenCL language version attributes from CIRGen. Preserve the compatible OpenCL version and the C++ for OpenCL version separately so later lowering does not infer one from the other.
Assisted-by: Codex / GPT-5.6 Sol
[CIR][OpenCL] Add OpenCL language version module attributes
Add structured CIR module attributes for OpenCL and C++ for OpenCL language versions. Verify their module-level placement and version components so lowering can consume explicit source-language version state.
Assisted-by: Codex / GPT-5.6 Sol
MC: Move BinutilsVersion from TargetOptions to MCTargetOptions
BinutilsVersion has no codegen use and only used by MCAsmInfo to check
ELF assembler features.
Co-authored-by: Claude (claude-opus-4.8) <noreply at anthropic.com>
[libc++][NFC] Avoid empty namespace in `<__concepts/common_with.h>` (#221407)
...in pre-C++20 modes. This fixes complaining from clang-tidy checks in
CI.
[OpenMPOpt] Ask the runtime how many of a block's threads can be workers
The custom state machine gates a thread on InitCB < BlockHwSize - WarpSize,
reconstructing the number of worker threads from the block size on the
assumption that the main thread occupies a whole warp above them. The DeviceRTL
already computes that number, in mapping::getMaxTeamThreads(), and its own
generic state machine gates on it in shouldEnterStateMachine(). Export it as
__kmpc_get_max_team_threads() and call that instead, so the compiler's state
machine and the runtime's agree by construction rather than by arithmetic that
has to be kept in step with the launch geometry.
This is NFC here: getMaxTeamThreads() in generic mode is BlockSize - WarpSize,
the same three instructions folded into one call. It is not NFC for a toolchain
whose launch geometry differs. In ROCm, CGOpenMPRuntimeGPU starts a single extra
thread rather than a warp -- "Only one additional thread is started, not an
entire warp" -- so thread_limit(1024) on a 64-lane target launches 961 threads
and the runtime reports 960 workers, while the state machine's own arithmetic
says 961 - 64 = 897. The threads in between are in neither group: the state
machine returns immediately for them, and the parallel region still hands them
[12 lines not shown]
[OpenMPOpt] Look inside the callbacks the loop runtime functions are handed
The __kmpc_{distribute_,for_,distribute_for_}static_loop_* functions receive the
loop body as a callback, so a parallel region written inside that body is
reachable from the kernel through the runtime call. AAKernelInfo could not see
that, and recorded the call as reaching an unknown parallel region. A kernel
using these functions therefore always got a worker state machine whose only
option was to indirectly call whatever work function it was handed.
Describe the callback argument of each of these functions in OMPKinds.def and
attach the corresponding !callback metadata in OpenMPOpt, then fold the
callback's AAKernelInfo state into the caller's. The state machine can now
dispatch directly to the regions the loop body actually reaches. Relax the two
"more than one callee means give up" checks for functions carrying !callback,
since the callback edge is a second edge by construction and is analyzable.
The conservative unknown-region record is kept for the case that motivated it, a
callback we only see a declaration of.
[39 lines not shown]
WebAssembly: Introduce ExceptionHandling::EmscriptenEH model
Add a dedicated EmscriptenEH exception model so the control uses
the standard exception model control, instead of relying on a backend
specific cl::opt. This will later migrate to a module flag and
remove -enable-emscripten-cxx-exceptions
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[LoongArch][RISCV] Ignore debug uses when merging base offsets (#221117)
## Summary
Debug instructions are currently enumerated as ordinary users by the RISC-V
and LoongArch merge-base-offset passes. A `DBG_VALUE` can therefore veto an
otherwise valid fold and make `-g` add an ordinary address-calculation
instruction.
Use the non-debug instruction iterator for validation and rewriting, and make
affected debug values unavailable before changing the address represented by
the destination register. Also handle the case where the register has no
ordinary users, which becomes possible after debug uses are excluded.
The same change is applied to both targets because their implementations and
failure mode are equivalent.
## Testing
[6 lines not shown]
DAGCombiner: Drop AllowFPOpFusion from visitFADDForFMACombine
Rewrites fp-dp3.ll to use flags on individual patterns. It weirdly
used different triples for the fp-contract on and off cases, seemingly
an artifact of the ARM64 and AArch64 merge.
fp-contract.cu is essentially a bugfix, the local fp contract(on) pragma
wins over the global flag now.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[X86] LowerCLMUL - improve vXi32 codegen (#221289)
Shuffle combining is struggling to handle the mixture of vector
unrolling, truncations and optimal use of PCLMULQDQ swizzles, this patch
gives us more optimal lowering direct instead of relying on fixup.
Use the PCLMULQDQ lo/hi control immediate to handle anyext i64 element
evaluation - along with shifting the upper elements down to the lowest
bits of the i64 in parallel
Use UNPACK build vector pattern to avoid FPR<->GPR traffic
CLMULH needs to be handled in a future patch, and vXi16 /might/ be worth
handling as well.
[Clang] Fix Crash in Sema::DiagnoseUnguardedAvailability On 'if' With No Condition (#220004)
**Problem**
C++ 23 introduced consteval expressions which allows `if` statements to
have no condition:
```
if consteval {
}
```
`DiagnoseUnguardedAvailability::TraverseIfStmt` Assumed `If->getCond()`
would never return a `nullptr`, causing a `nullptr` dereference.
Fixes #219948