[mlir][Vector] Reject scalable reduction in warp distribution (#225267)
`WarpOpReduction` distributes a `vector.reduction` across warp lanes by
computing `numElements = vectorType.getShape()[0] / warpSize` and
building the per-lane type as a plain `VectorType::get({numElements},
...)`, with no check for a scalable operand.
For a scalable reduction vector, e.g. `vector<[32]xf32>` with warp size
32, this silently drops the scalable marker: the per-lane type becomes
the fixed `vector<1xf32>` instead of `vector<[1]xf32>`, so only 32 total
elements get reduced across lanes instead of `32*vscale`.
Reject a scalable reduction operand before any IR is created, and add a
negative test.
NVPTX: Drop LiveVariables from the register allocation pipeline (#225183)
The optimized RegAlloc pipeline ran LiveVariables only to satisfy
PHIElimination and TwoAddressInstruction, both of which no longer need it.
Remove the LiveVariables run (and, in the new pass manager, the
UnreachableMachineBlockElim that was there only as a LiveVariables
prerequisite).
Co-authored-by: Claude (Claude-Opus-4.8)
[LAA] Add tests with loops with may-not-return calls (NFC) (#228102)
Add tests for runtime-check bounds of loops containing a call that may
not return.
[clang-repl] Flush CUDA device bootstrap module before the first PTU (#226975)
The host path sets the bootstrap module aside with CacheCodeGenModule()
once the initial action has run, so the first PTU starts from a fresh
module. The CUDA device path skipped this step. The device module that
HandleTranslationUnit had already finalized during bootstrap stayed
current, and the first device PTU finalized it a second time, resulting
in CodeGen adding every module flag twice. The IR verifier then rejects
the module.
To reproduce, on a build with assertions, or with
`-fverify-intermediate-code` (hidden on release builds since the driver
disables the verifier there), the first input to `clang-repl --cuda`
fails:
```
module flag identifiers must be unique (or of 'require' type)
!"nvvm-reflect-ftz"
module flag identifiers must be unique (or of 'require' type)
[12 lines not shown]
CodeGen: Drop the LiveVariables parameter from convertToThreeAddress (#225182)
This was used for analysis updates, but now the analysis is being removed.
Co-authored-by: Claude (Claude-Opus-4.8)
[lldb][Windows] Move extended info to host thread (#227470)
I want to add support for reading thread locals on Windows. To do this,
we need to know the address of the TEB. This was implemented for
`TargetThreadWindows`, but when using lldb-server, we didn't have this
info. As a first step, move this to the `HostThreadWindows`, so both the
target thread and native thread can call it.
[DataFlowSanitizer] Properly add ext attributes on arguments as needed. (#225443)
TargetLibraryInfo is used to compute the extension attributes, in part by a new
getExtAttrForI8Param() method. It currently always returns ZExt (or SExt) but is
used so that a target can easily override this if needed.
[CodeGen] Remove unused SpillPlacement::Linked (NFC) (#227989)
The last use was removed on May 19, 2016 in commit
b926bdac4c18e0f31d827dec482f207856e88e1e.
Assisted-by: Antigravity
[OpenMP] Fix missing implicit barriers in device worksharing loops (#227735)
Device worksharing lowering discarded `NeedsBarrier`, omitting implicit
barriers after loops without `nowait`. Forward the requirement and emit
the barrier outside the outlined loop body, ensuring all participating
threads synchronise.
Co-authored-by: Codex <codex at openai.com>
[LAA] Add stencil group merging to reduce runtime pointer checks (#187252)
Take this loop, where S is only known at runtime:
for (i = 0; i < N; i++)
Out[i] = In[i - S] + In[i - 1] + In[i] + In[i + 1] + In[i + S];
groupChecks puts In[i - 1], In[i] and In[i + 1] in one group, because
their bounds differ by a constant. In[i - S] and In[i + S] differ from
the rest by a multiple of S, so each stays in its own group. That is
3 groups for In and 3 checks against Out. With two or three strides,
as in 3D stencils, the count grows fast. The Einstein Toolkit / Cactus
CCZ4 code has loops with thousands of checks, and the vectorizer gives
up on them.
This patch adds mergeStencilGroups, which runs after groupChecks. For
the loop above it makes one group for In, from In[i - S] to In[i + S].
One check against Out is enough, plus a check that 1 <= S <= Max.
[27 lines not shown]
[Offload][L0][NFC] Remove old ELF format support (#228011)
After we switched to using Offload Binary (for OpenMP) or direct SPIR-V
images (for SYCL) this support is not used anymore.
Assisted by Claude.
[TableGen] Allow AsmWriter to generate uint64_t tables. NFC. (#227641)
Each OpInfo entry in the generated AMDGPUInstPrinter::getMnemonic
carries 8 bytes of data. Previously GenAsmWriter would split that into
two uint32_t tables for no good reason. Generating a single uint64_t
table makes for shorter output and slightly better generated code.
[AMDGPU][SROA] Expand cast chain handling to floating point types
CreateBitPreservingCastChain does not produce inttoptr or ptrtoint for floating-point types. When promoting structs like { float, float } to <2 x float>, this can lead to pointers being bitcast directory to <2 x float> which is invalid. This change expands the use of the intermediate to these cases
[clang][CIR] Route .cir cc1 input to the ClangIR frontend action (#227156)
Add `FrontendAction::hasCIRSupport()` and extend the IR bypass in
BeginSourceFile to `Language::CIR`. CIRGenAction opts in and, for now,
reports that ClangIR input is not yet supported; parsing and lowering
follow in separate patches. Actions without CIR support report
`err_ast_action_on_cir`. ClangIR input implies `-fclangir*`, since no
other pipeline can consume it.
This is part of the experimental .cir round-trip capability discussed in
https://discourse.llvm.org/t/90998; the input format carries no
stability guarantee.
\* Happy to emit diagnostics here if we want to be more defensive on
this.
---------
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[CIR] Make ASTContext optional in runCIRToCIRPasses (#227128)
Read the triple and the new cir.target_abi attribute from the module
instead of the ASTContext, erroring if either is missing, so the
CIR-to-CIR pipeline can run without a live AST. Part of preparing for:
https://discourse.llvm.org/t/rfc-clangir-making-cir-pipeline-boundaries-first-class-driver-artifacts/90998/15
---------
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[SDPatternMatch] Simplify MaxMin_match. NFC (#227947)
Remove EffectiveOperands and m_SpecificOpc since we no longer need to
match VP nodes.
Use . instead of -> for calling getOperand, getOpcode, and
getNumOperands.
[InstCombine] Fix profile propagation in uordered-fcmp-select.ll (#228090)
The condition simply gets inverted so we can directly propagate the
branch weights and then swap them afterwards.
[CIR][CodeGen][NFC] Share hasExtraNeonArgument
Deduplicates `hasExtraNeonArgument` between CIR and classic CodeGen into
`TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share the Arm SME inlinability check
Deduplicates `ArmSMEInlinability` and `getArmSMEInlinability` between CIR and
classic CodeGen into a new `TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).