AMDGPU: Remove amdgpu-scalarize-global-loads option (#230744)
This was added to avoid updating many tests when the global load to scalar
load optimization was first added, but all tests have now converted off of it.
This was also weirdly wired into the subtarget.
[clang-tidy] Improve modernize-use-equals-default check to handle explicit parent constructor calls (#226456)
Now capture this pattern:
struct Base {};
struct C : Base {
C() : Base() {}
};
And turn it into
struct Base {};
struct C : Base {
C() = default;
};
AMDGPU: Remove amdgpu-scalarize-global-loads option
This was added to avoid updating many tests when the global load to scalar load
optimization was first added, but all tests have now converted off of it. This
was also weirdly wired into the subtarget.
Revert "[KnownFPClass] Refine positive zero result for sqrt" (#230741)
Reverts llvm/llvm-project#214987
Causes a miscompile with nsz: `fcmp une (sqrt nsz fp128 -0.0), 0.0`
folds to true. Reported by Linaro CI (Fujitsu C/0055/0055_0004,
LLVM-2292).
https://godbolt.org/z/cYb3Gah6n
[Instrumentation] Stop exporting cl::opts. NFC (#230725)
Replace the `extern cl::opt` declarations of -profile-correlate in
FrontendDriver and -pgo-instrument-cold-function-only in Passes with
getters, and delete clang's unused declaration of the former. This
prepares for moving Instrumentation's options into TableGen.
Aided by Opus 5.5
[MLGO] Remove std includes from pre-generated test models (#230724)
c90b6414e607 (#227941) always registers the test models in
llvm/lib/Analysis/models/*.inc.
The generated InlinerModels.h and RegAllocEvictModels.h include each
model header inside a namespace.
With LLVM_ENABLE_MODULES=ON, the models' `#include <map>` and `#include
<string>` become module imports inside that namespace, which clang
rejects:
error: redundant #include of module 'std_map' appears within namespace
'llvm::regalloc_Model1_ns'
[-Wmodules-import-nested-redundant]
This patch removes the includes, since MLInlineAdvisor.cpp and
MLRegAllocEvictAdvisor.cpp already include <map> and <string> before the
model headers.
[3 lines not shown]
[Mips] Remove unused -mips-fix-global-base-reg (#230733)
Unused since 3ecc5273c148 (2012). Remove its RUN
lines from tls-static.ll, whose STATIC32/STATIC64 checks cover the same
output, and delete the disabled global-pointer-reg.ll.
[BOLT] Restore -relax-plt as a deprecated no-op option
c6ee71d5e8a5 removed -relax-plt from the AArch64 relaxation pass, which
makes llvm-bolt reject existing command lines that still pass it.
Re-add the option as a hidden no-op that only prints a deprecation
warning, so scripts using it keep working.
[RISCV] Declare command line options in TableGen (#230014)
Move the cl::opts of RISCVCodeGen into RISCVOptions.td, and those of
RISCVDesc and RISCVAsmParser into MCTargetDesc/RISCVMCOptions.td,
registered by LLVMInitializeRISCVTargetMC(). RISCVTargetMachine holds
`const RISCVOptions &CLOpts` and RISCVSubtarget copies it;
RISCVAsmBackend, RISCVInstPrinter, and RISCVTargetStreamer hold `const
RISCVMCOptions &CLOpts`. -riscv-rvv-regalloc (RegisterPassParser) stays
cl::opt.
`llvm-objdump -M emit-x8-as-fp` now sets a file-static flag instead of
writing the cl::opt.
-riscv-br-merging-base-cost now takes effect when given once; it
previously required getNumOccurrences() > 1.
Aided by Opus 5.5
[AMDGPU] Fold add of a variable into a zero dot accumulator
When the dot intrinsic has a zero accumulator, no clamp and a single
add user, fold the add operand into the accumulator:
```llvm
%dot = call i32 @llvm.amdgcn.sdot4(i32 %a, i32 %b, i32 0, i1 false)
%r = add i32 %dot, %x
=>
%r = call i32 @llvm.amdgcn.sdot4(i32 %a, i32 %b, i32 %x, i1 false)
```
If %x is defined after the dot in the same block, the dot is moved
down to the add. The fold is skipped across blocks to avoid sinking
the dot into a loop.
[Clang][AMDGPU] Use unsigned int for tensor builtin D# groups (#230714)
The D# tensor descriptor groups of __builtin_amdgcn_tensor_load_to_lds
and __builtin_amdgcn_tensor_store_from_lds are bit fields, not signed
values. D0 was already declared as a vector of unsigned int; D1 through
D4 are now consistent with it.
OpenCL does not allow lax vector conversions, so the tests are updated
to pass unsigned vectors. The generated IR is unchanged.
Reference: https://github.com/llvm/llvm-project/pull/193310
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[lldb][docs] Use project-local documentation links
Replace same-project absolute URLs with relative Markdown links and Sphinx
cross-references so local and archived documentation stays self-contained.
Update generated Python API docstrings accordingly, and enable the absolute
link check for the LLDB docs to keep it that way.
[docs] Enable absolute documentation link checks
Configure the LLVM documentation URL prefixes so the Sphinx build rejects
absolute links to documents in the same project. Clang is already configured.
Part of #214861
[clang][Parse] Delay template-id destruction in NTTP default arguments (#230513)
There was a UAF when the lambda appears within NTTP default arguments.
Co-authored-by: Emery Conrad <emery.conrad at chicagotrading.com>
Co-authored-by: Sebastian Schwartz <sebastian.schwartz at chicagotrading.com>
Co-authored-by: Emery Conrad <emery.conrad at chicagotrading.com>
Assisted-by: Claude Code (claude-opus-5-5)
[CIR] Add bytecode encodings for attributes (#229586)
Without a BytecodeDialectInterface, MLIR encodes a dialect's attributes
through their assembly format, embedding the printed text in the
bytecode. This adds native encodings for 27 CIR attributes (constants,
constant initializers, global views, method/data-member pointers, C++
special-member attributes, and small metadata attributes), modelled on
the LLVM dialect's bytecode support.
On attribute-dense modules this shrinks the bytecode ~30-40% (8.6KB to
5.4KB across the new tests). The room ahead is huge, emitting bytecode
for CIRGenModule.cpp (441MB of CIR text) still exceeds 22GB RSS and 30
minutes with these encodings (64GB and 45 minutes without), since types
keep using the assembly fallback. This is paving towards selfhosting
with bytecode.
Coverage is partial by design: anything not listed keeps using the
assembly-format fallback, exactly as before. Types keep the fallback too
until the type encodings land separately.
[3 lines not shown]
[scudo] Check quarantine batch bounds in release builds (#230247)
[Attacking Scudo's Quarantine, section 3: Write Where
Ptr](https://un1fuzz.github.io/articles/quarantine_attack.html#a3)
describes corrupting a quarantine batch's `Count` so that enqueue writes
the freed pointer outside the batch. The original [PoC and exploit code
are
here](https://github.com/un1fuzz/scudo_research/tree/main/quarantine_arbitrary_return).
`push_back()` currently guards its index with a debug-only check; an
oversized count also bypasses enqueue's full-batch equality check.
Make that bound a release-build check. Also validate both counts before
deciding whether batches can merge, enforce merge capacity in
production, and check the count before shuffling (which precedes
recycling). The capacity comparison uses subtraction after validating
both operands, avoiding corrupted-count addition wrapping around.
This closes the out-of-bounds enqueue primitive described in section 3.
It does not address section 2's Double Return attack using in-range
[23 lines not shown]
[CIR][NFC] Add missing NYI handling for some x86 builtins (#230609)
While doing a recent code review, I noticed that there were some x86
builtins that were incorrectly falling through to code that handles
builtins below them in a switch. This change adds an errorNYI diagnostic
rather than falling through.
AMDGPU: Index loads by workitem id in tests shared with r600 (#230667)
These tests relied on -amdgpu-scalarize-global-loads=false to select
vector loads from uniform pointer arguments. They also have r600 run
lines, so keep the kernels and index the input pointers by the workitem
id instead. Also fix shl_v2i16 not using its computed pointers, and
v_shl_32_i64 using the workgroup id. Also add some uniform variants
of some cases.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
fix(AMDGPU): guard OR folds with shared conditions
A shared uniform condition still needs materializing after folding a
divergent OR to a select, and sharing can introduce an extra SCC
conversion. The extension's single-use check does not prevent this
code-size regression.
Check the condition's uses and divergence as well. Add regression and
control cases, and consolidate the boolean OR tests in or.ll.