TableGen: Cache resolved class instantiations
The same class instantiation is often referenced many times in one
record, and each reference was resolved again from scratch along with
everything nested in its arguments. Cache the results in MapResolver and
RecordResolver. The cache is cleared when the mappings change, and is
bypassed while a mapped variable is being resolved.
Nested instantiations are resolved through a TrackUnresolvedResolver,
which does not change the result. These share the cache of the
underlying resolver, which also records whether any references were left
unresolved so the trackers can be updated on a hit.
Gigainstructions executed by llvm-tblgen --null-backend
before after
AMDGPU.td 71.75 18.22 -74.6%
X86.td 6.10 6.15 +0.8%
AArch64.td 3.28 3.32 +1.3%
[3 lines not shown]
TableGen: Skip resolving values that cannot change
Lists, dags, bits and arguments made only of leaf values always resolve
to themselves, but resolving them still walked every element. Record
this when they are created and return early.
Gigainstructions executed by llvm-tblgen --null-backend
before after
AMDGPU.td 79.12 71.73 -9.3%
X86.td 6.25 6.12 -2.1%
AArch64.td 3.40 3.28 -3.7%
RISCV.td 8.34 8.19 -1.8%
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
TableGen: Use DenseMap for the IntInit pool
Gigainstructions executed by llvm-tblgen --null-backend
before after
AMDGPU.td 79.34 79.13 -0.3%
X86.td 6.30 6.25 -0.7%
AArch64.td 3.43 3.41 -0.6%
RISCV.td 8.40 8.34 -0.7%
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[VPlan] Use getFastMathFlagsOrNone in VPIRFlags::applyFlags (NFC) (#231036)
applyFlags hand-rolled the same field-by-field FastMathFlagsTy to
FastMathFlags conversion as getFastMathFlagsOrNone. Use the latter
together with copyFastMathFlags.
AMDGPU: Use TargetSubtarget.td in TargetDef tablegen tests (#231018)
The -gen-amdgpu-target-def tests only need the subtarget and processor
classes. Include TargetSubtarget.td instead of Target.td, which also
checks that it works standalone, and drop the unneeded dummy Target.
The 14 llvm-tblgen RUN lines each parsed Target.td, which costs 1.29 G
instructions:u (Release+asserts); TargetSubtarget.td costs 0.005 G.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[gn] port c90b6414e607 (mlgo EmitC models) (#231035)
Ports ab547095ead5 since it's needed unconditionally after c90b6414e607.
Register only the checked-in, pre-lowered _test_inc test model.
Lowering .mlir models is still not implemented in the GN build.
[gn] make write_file.py replace literal \n with newlines (#231034)
Like write_cmake_config.py. This makes it possible to write multi-line
generated headers from an action(), which is needed for the next commit.
As demo, update the only current existing caller to write a
trailing newline. This matches the CMake build.
[AMDGPU] Limit the fmul fusion discount to a matching context type
The fmul is free when its context instruction feeds a fusable fadd or
fsub. A vector fmul priced with a scalar lane as the context inherits
that fusion only when it is emitted lane by lane. On a packed type the
scalar user does not show that the vector fmul feeds a vector fadd, and
products extracted into a scalar fadd chain keep the packed fmul while
the fma is lost.
[llubi] Use the correct poison representation for byte values (#231028)
We always use `ByteValue` to represent `bN poison`.
The test is generated by DeepSeek-V4.1-Flash. Looks like there is no
other bug like this.
RegAllocGreedy: Avoid unused block frequency lookups in hint recoloring (#231027)
collectHintInfo looked up the block frequency of every copy, but the
frequencies are only used for the profitability check when the register
would be recolored. Most visited registers are already assigned the
target register, so record the block and look the frequency up only
when it is needed. Also use a SmallDenseSet for the visited set, which
otherwise falls back to a std::set on anything but tiny components.
On a stress test with 4000 critical edges into a block with 4 PHIs,
this roughly halves the time in Greedy.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[RISCV][P-ext] Change RISCVISD::SATI to use [1,bitwidth] for the immediate. (#230377)
This matches the intrinsics and assembly syntax for the instruction.
Assisted-by: Claude
AMDGPU: Use TargetSubtarget.td in TargetDef tablegen tests
The -gen-amdgpu-target-def tests only need the subtarget and processor
classes. Include TargetSubtarget.td instead of Target.td, which also
checks that it works standalone, and drop the unneeded dummy Target.
The 14 llvm-tblgen RUN lines each parsed Target.td, which costs 1.29 G
instructions:u (Release+asserts); TargetSubtarget.td costs 0.005 G.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
AMDGPU: Use minimal tablegen inputs for TargetParser (#231017)
Avoid parsing the heavyweight AMDGPU.td to generate the TargetParser
inc files. Only include the minimal subtarget information required.
llvm-min-tblgen -gen-amdgpu-target-def (Release+asserts):
wall instructions:u peak RSS .td deps
AMDGPU.td 12.16 s 79.49 G 909 MB 76
AMDGPUTargetParserDef.td 0.01 s 0.056 G 7 MB 10
R600.td 0.38 s 1.59 G 92 MB 57
R600TargetParserDef.td <0.01 s 0.008 G 6 MB 10
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
RegAllocGreedy: Avoid unused block frequency lookups in hint recoloring
collectHintInfo looked up the block frequency of every copy, but the
frequencies are only used for the profitability check when the register
would be recolored. Most visited registers are already assigned the
target register, so record the block and look the frequency up only
when it is needed. Also use a SmallDenseSet for the visited set, which
otherwise falls back to a std::set on anything but tiny components.
On a stress test with 4000 critical edges into a block with 4 PHIs,
this roughly halves the time in Greedy.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
Attributor: Use NullPointerIsDefined for null nocapture check (#229333)
Previously this did not try to respect the null_pointer_is_valid attribute.
Query NullPointerIsDefined with the anchor scope, matching the handling
in AANoAlias::isImpliedByIR.
Note: a valid null passed to a capturing callee still gets noalias at the
call site. AANoAliasCallSiteArgument queries IRPosition::value() of the
constant, which has no anchor scope, so null_pointer_is_valid is not
consulted there.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
AMDGPU: Use TargetSubtarget.td in TargetDef tablegen tests
The -gen-amdgpu-target-def tests only need the subtarget and processor
classes. Include TargetSubtarget.td instead of Target.td, which also
checks that it works standalone, and drop the unneeded dummy Target.
The 14 llvm-tblgen RUN lines each parsed Target.td, which costs 1.29 G
instructions:u (Release+asserts); TargetSubtarget.td costs 0.005 G.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
AMDGPU: Use minimal tablegen inputs for TargetParser
Avoid parsing the heavyweight AMDGPU.td to generate the TargetParser
inc files. Only include the minimal subtarget information required.
llvm-min-tblgen -gen-amdgpu-target-def (Release+asserts):
wall instructions:u peak RSS .td deps
AMDGPU.td 12.16 s 79.49 G 909 MB 76
AMDGPUTargetParserDef.td 0.01 s 0.056 G 7 MB 10
R600.td 0.38 s 1.59 G 92 MB 57
R600TargetParserDef.td <0.01 s 0.008 G 6 MB 10
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
TableGen: Split subtarget and processor classes out of Target.td (#231016)
Move the subtarget feature, predicate and processor definitions into a
separate file. In the future this will allow TargetParser tablegen inputs to
use them without parsing the full target definitions like instructions,
registers and intrinsics.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[CIR] Add missing Coroutine region successors (#230466)
* A coroutine may be destroyed after its `initial_suspend`. Add an edge
from there to `destroy`
* `suspend` has no edge to `resume`. Add this, too
[Option] Add Condition for build-conditional library options (#230898)
A field under `let Condition = AssertsOnly` declares its member, table
row, and parser case under `#if !defined(NDEBUG)`. In other builds, the
command line rejects the option as an unknown argument, and readers must
be guarded like those of a conditional cl::opt. An OptionsStruct's table
rows refer to groups and alias targets by OPT_ enumerator, since an
omitted row shifts the IDs after it.
Port -loop-fusion-verbose-debug and -stress-ivchain, the NDEBUG-only
Scalar options.
Aided by Opus 5.5
[WebKit checkers] Treat destruction of a null smart pointer temporary as trivial (#230809)
TrivialFunctionAnalysis treated HashMap<K, Ref<V>>::get as non-trivial
because its miss path, MappedTraits::peek(MappedTraits::emptyValue()),
creates and destroys a temporary Ref<V> constructed from
HashTableEmptyValue. Ref's destructor calls deref(), which may delete,
so any local raw pointer obtained from such a HashMap was reported by
alpha.webkit.UncountedLocalVarsChecker even in otherwise trivial
contexts.
Skip the destructor check of a temporary owning smart pointer when it is
known to hold nullptr: one which is default constructed, constructed
from nullptr or HashTableEmptyValue, copied or moved from such a smart
pointer, or returned by a function whose body is a single return
statement of such a smart pointer (e.g.
HashTraits<Ref<T>>::emptyValue()).
Also treat CXXScalarValueInitExpr (e.g. T() for a scalar T as in
GenericHashTraits<T>::emptyValue()) as trivial.
Authored with Claude Code.
[RISCV][P-ext] Re-generate rvp-intrinsics.c. NFC (#230691)
This fixes cases that used RV32 and RV64 instead of the common CHECK.
This also fixes cases where a mix of CHECK and RV32/RV64 were used in
the same function. The script is unable to do this so they must have
been manually written. Manual checks place burden on future editors of
this file since they can't use the script output directly.
[RISCV][P-ext] Move scalar mulh intrinsic tests to rv32p.ll and rv64p.ll. NFC (#230716)
These were in a file with simd in its name, but they aren't simd.
Assisted-by: Claude