[AMDGPU] Limit the fmul fusion discount to a matching context type
The fmul is free when its context instruction feeds a fusable fadd or
fsub. A vector fmul priced with a scalar lane as the context inherits
that fusion only when it is emitted lane by lane. On a packed type the
scalar user does not show that the vector fmul feeds a vector fadd, and
products extracted into a scalar fadd chain keep the packed fmul while
the fma is lost.
[llubi] Use the correct poison representation for byte values (#231028)
We always use `ByteValue` to represent `bN poison`.
The test is generated by DeepSeek-V4.1-Flash. Looks like there is no
other bug like this.
RegAllocGreedy: Avoid unused block frequency lookups in hint recoloring (#231027)
collectHintInfo looked up the block frequency of every copy, but the
frequencies are only used for the profitability check when the register
would be recolored. Most visited registers are already assigned the
target register, so record the block and look the frequency up only
when it is needed. Also use a SmallDenseSet for the visited set, which
otherwise falls back to a std::set on anything but tiny components.
On a stress test with 4000 critical edges into a block with 4 PHIs,
this roughly halves the time in Greedy.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[RISCV][P-ext] Change RISCVISD::SATI to use [1,bitwidth] for the immediate. (#230377)
This matches the intrinsics and assembly syntax for the instruction.
Assisted-by: Claude
AMDGPU: Use TargetSubtarget.td in TargetDef tablegen tests
The -gen-amdgpu-target-def tests only need the subtarget and processor
classes. Include TargetSubtarget.td instead of Target.td, which also
checks that it works standalone, and drop the unneeded dummy Target.
The 14 llvm-tblgen RUN lines each parsed Target.td, which costs 1.29 G
instructions:u (Release+asserts); TargetSubtarget.td costs 0.005 G.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
AMDGPU: Use minimal tablegen inputs for TargetParser (#231017)
Avoid parsing the heavyweight AMDGPU.td to generate the TargetParser
inc files. Only include the minimal subtarget information required.
llvm-min-tblgen -gen-amdgpu-target-def (Release+asserts):
wall instructions:u peak RSS .td deps
AMDGPU.td 12.16 s 79.49 G 909 MB 76
AMDGPUTargetParserDef.td 0.01 s 0.056 G 7 MB 10
R600.td 0.38 s 1.59 G 92 MB 57
R600TargetParserDef.td <0.01 s 0.008 G 6 MB 10
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
RegAllocGreedy: Avoid unused block frequency lookups in hint recoloring
collectHintInfo looked up the block frequency of every copy, but the
frequencies are only used for the profitability check when the register
would be recolored. Most visited registers are already assigned the
target register, so record the block and look the frequency up only
when it is needed. Also use a SmallDenseSet for the visited set, which
otherwise falls back to a std::set on anything but tiny components.
On a stress test with 4000 critical edges into a block with 4 PHIs,
this roughly halves the time in Greedy.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
Attributor: Use NullPointerIsDefined for null nocapture check (#229333)
Previously this did not try to respect the null_pointer_is_valid attribute.
Query NullPointerIsDefined with the anchor scope, matching the handling
in AANoAlias::isImpliedByIR.
Note: a valid null passed to a capturing callee still gets noalias at the
call site. AANoAliasCallSiteArgument queries IRPosition::value() of the
constant, which has no anchor scope, so null_pointer_is_valid is not
consulted there.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
AMDGPU: Use TargetSubtarget.td in TargetDef tablegen tests
The -gen-amdgpu-target-def tests only need the subtarget and processor
classes. Include TargetSubtarget.td instead of Target.td, which also
checks that it works standalone, and drop the unneeded dummy Target.
The 14 llvm-tblgen RUN lines each parsed Target.td, which costs 1.29 G
instructions:u (Release+asserts); TargetSubtarget.td costs 0.005 G.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
AMDGPU: Use minimal tablegen inputs for TargetParser
Avoid parsing the heavyweight AMDGPU.td to generate the TargetParser
inc files. Only include the minimal subtarget information required.
llvm-min-tblgen -gen-amdgpu-target-def (Release+asserts):
wall instructions:u peak RSS .td deps
AMDGPU.td 12.16 s 79.49 G 909 MB 76
AMDGPUTargetParserDef.td 0.01 s 0.056 G 7 MB 10
R600.td 0.38 s 1.59 G 92 MB 57
R600TargetParserDef.td <0.01 s 0.008 G 6 MB 10
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
TableGen: Split subtarget and processor classes out of Target.td (#231016)
Move the subtarget feature, predicate and processor definitions into a
separate file. In the future this will allow TargetParser tablegen inputs to
use them without parsing the full target definitions like instructions,
registers and intrinsics.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[CIR] Add missing Coroutine region successors (#230466)
* A coroutine may be destroyed after its `initial_suspend`. Add an edge
from there to `destroy`
* `suspend` has no edge to `resume`. Add this, too
[Option] Add Condition for build-conditional library options (#230898)
A field under `let Condition = AssertsOnly` declares its member, table
row, and parser case under `#if !defined(NDEBUG)`. In other builds, the
command line rejects the option as an unknown argument, and readers must
be guarded like those of a conditional cl::opt. An OptionsStruct's table
rows refer to groups and alias targets by OPT_ enumerator, since an
omitted row shifts the IDs after it.
Port -loop-fusion-verbose-debug and -stress-ivchain, the NDEBUG-only
Scalar options.
Aided by Opus 5.5
[WebKit checkers] Treat destruction of a null smart pointer temporary as trivial (#230809)
TrivialFunctionAnalysis treated HashMap<K, Ref<V>>::get as non-trivial
because its miss path, MappedTraits::peek(MappedTraits::emptyValue()),
creates and destroys a temporary Ref<V> constructed from
HashTableEmptyValue. Ref's destructor calls deref(), which may delete,
so any local raw pointer obtained from such a HashMap was reported by
alpha.webkit.UncountedLocalVarsChecker even in otherwise trivial
contexts.
Skip the destructor check of a temporary owning smart pointer when it is
known to hold nullptr: one which is default constructed, constructed
from nullptr or HashTableEmptyValue, copied or moved from such a smart
pointer, or returned by a function whose body is a single return
statement of such a smart pointer (e.g.
HashTraits<Ref<T>>::emptyValue()).
Also treat CXXScalarValueInitExpr (e.g. T() for a scalar T as in
GenericHashTraits<T>::emptyValue()) as trivial.
Authored with Claude Code.
[RISCV][P-ext] Re-generate rvp-intrinsics.c. NFC (#230691)
This fixes cases that used RV32 and RV64 instead of the common CHECK.
This also fixes cases where a mix of CHECK and RV32/RV64 were used in
the same function. The script is unable to do this so they must have
been manually written. Manual checks place burden on future editors of
this file since they can't use the script output directly.
[RISCV][P-ext] Move scalar mulh intrinsic tests to rv32p.ll and rv64p.ll. NFC (#230716)
These were in a file with simd in its name, but they aren't simd.
Assisted-by: Claude
AMDGPU: Use TargetSubtarget.td in TargetDef tablegen tests
The -gen-amdgpu-target-def tests only need the subtarget and processor
classes. Include TargetSubtarget.td instead of Target.td, which also
checks that it works standalone, and drop the unneeded dummy Target.
The 14 llvm-tblgen RUN lines each parsed Target.td, which costs 1.29 G
instructions:u (Release+asserts); TargetSubtarget.td costs 0.005 G.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
AMDGPU: Use minimal tablegen inputs for TargetParser
Avoid parsing the heavyweight AMDGPU.td to generate the TargetParser
inc files. Only include the minimal subtarget information required.
llvm-min-tblgen -gen-amdgpu-target-def (Release+asserts):
wall instructions:u peak RSS .td deps
AMDGPU.td 12.16 s 79.49 G 909 MB 76
AMDGPUTargetParserDef.td 0.01 s 0.056 G 7 MB 10
R600.td 0.38 s 1.59 G 92 MB 57
R600TargetParserDef.td <0.01 s 0.008 G 6 MB 10
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
TableGen: Split subtarget and processor classes out of Target.td
Move the subtarget feature, predicate and processor definitions into
a separate file. In the future this will allow TargetParser tablegen
inputs to use them without parsing the full target definitions like
instructions, registers and intrinsics.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
Fix GCC warning by enclosing operand with parentheses. (#230939)
PR #174549 made a change to `APValue.h` which causes GCC to emit a
warning like this:
```
clang/include/clang/AST/APValue.h: In static member function ‘static clang::DynamicAllocLValue clang::DynamicAllocLValue::getFromOpaqueValue(const void*)’:
clang/include/clang/AST/APValue.h:114:54: warning: suggest parentheses around ‘-’ in operand of ‘&’ [-Wparentheses]
114 | V.AllocKind = Combined & (1 << NumAllocKindBits) - 1;
| ~~~~~~~~~~~~~~~~~~~~~~~~^~~
```
From my build with gcc, this shows up around 1200 times in the logfile
for my build. The problem with this is that for distributed builds, this
forces the build of the file where the warning is emitted to be done
locally which causes a bottleneck. This fixes the source of the warning
so that the build of the files which include `APValue.h` can again be
distributed.
[AMDGPU] Extract byte lanes of a split vector from the 32-bit source (#228424)
After a <4 x i8> is split into i16 halves, a byte lane is extended
from an i16 shift of a truncate. Rewrite it on the 32-bit source as an
and of srl or a sign_extend_inreg of srl, which select to a single bit
field extract.
Uniform zero and any extends are left alone, their i16 shift is already
promoted to i32.
Assisted-by: Claude Code Opus 5
AMDGPU: Move GCN subtarget features to GCNFeatures.td (#230996)
Move the GCN SubtargetFeature definitions, the FeatureISAVersion lists
and AMDGPUFrontendVisibleFeatures out of AMDGPU.td, so they can be used
without parsing the instruction, register and intrinsic definitions.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
---
<sub>Stack created with <a
href="https://github.com/github/gh-stack">GitHub Stacks CLI</a> • <a
href="https://gh.io/stacks-feedback">Give Feedback 💬</a></sub>
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
AMDGPU: Move R600 subtarget features to R600Features.td (#230997)
Split the R600 SubtargetFeature definitions and
R600FrontendVisibleFeatures out of R600Processors.td, mirroring
GCNFeatures.td.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
---
<sub>Stack created with <a
href="https://github.com/github/gh-stack">GitHub Stacks CLI</a> • <a
href="https://gh.io/stacks-feedback">Give Feedback 💬</a></sub>
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[libc] Implement clearenv (#230941)
Added the clearenv entrypoint to clear all environment variables.
* Declared clearenv in stdlib.yaml under GNU standards.
* Added EnvironmentManager::clear() to reset managed storage and set
environ to NULL, leaving string buffers intact to prevent use-after-free
for callers holding getenv() results.
* Added the entrypoint for Linux targets and enabled it for full builds
on aarch64, riscv, and x86_64.
* Added integration test covering clearenv, subsequent environment
modifications, idempotence, and pointer retention.
Assisted-by: Automated tooling, human reviewed.