[SCEV] - Add positive-stride predicate for backedge-taken count. (#222261)
When `howManyLessThans` encounters a loop with an unknown stride that
cannot be proven finite (no `mustprogress` or side-effect-free
guarantee), SCEV currently returns `CouldNotCompute` for the
backedge-taken count. This blocks downstream consumers like the loop
vectorizer from optimizing such loops.
This patch relaxes the requirement by allowing a predicated
backedge-taken count when `AllowPredicates` is true. Instead of
requiring `loopIsFiniteByAssumption(L)` unconditionally, we add a
`Compare predicate: stride sgt 0` when finiteness cannot be proven.
A positive stride guarantees forward progress, making the BTC formula
correct. The predicate is emitted as a runtime check by consumers
(e.g., the loop vectorizer generates a guard branch before the vector
loop).
[17 lines not shown]
[docs] Pass -c when generating a PCH file (#226850)
Without options like -fsyntax-only/-E/-S/-c, the driver runs in link
mode, and a header-only command line produces a PCH only incidentally: a
linker option such as -lm or -Wl,..., including one from a config file,
adds a link job. Use -c in the PCH examples.
Also fix two broken examples: -ignore-pch is a driver option (cc1
rejects -Xclang -ignore-pch), and the relocatable PCH example is
missing -o. In LibASTImporter.md, generate the C++ AST files with
-emit-ast instead of treating .cpp files as headers.
LLM-aided
[lldb][test] Handle None function name in frame provider tests (#226936)
Fixes #191859.
On Arm one of the functions has no name, which is a valid state.
```
Num frames: 7
0 frame #0: 0x092a151c a.out`baz at main.c:5:3 --name: baz
1 frame #1: 0x092a1540 a.out`bar at main.c:10:10 --name: bar
2 frame #2: 0x092a1560 a.out`foo at main.c:15:10 --name: foo
3 frame #3: 0x092a158c a.out`main at main.c:20:10 --name: main
4 frame #4: 0xf7ab939a libc.so.6` --name: None
5 frame #5: 0xf7ab943e libc.so.6`__libc_start_main + 94 --name: __libc_start_main
6 frame #6: 0x092a1434 a.out`_start + 40 --name: _start
```
The prefix provider needs to handle that possibility.
Though the Python type annotations on ScriptedFrame say that the name
methods return string, in reality they can return None, which the C++
sides expect. I will fix that in another PR.
[LLVMIR] Directly lower global initializer GEP to constant expression (#226904)
MLIR lowers global initializers in a very unusual way, by using an
IRBuilder without insertion point, and relying on the fact that for a
well-formed initializer, all the instructions will be converted to
constant expressions and no actual instructions that require insertion
will be produced.
This causes issues for https://github.com/llvm/llvm-project/pull/226425,
which requires an insertion point for CreateGEP to determine the data
layout.
To avoid this, make convertGEPOp() directly produce a constant
expression GEP if there is no insertion point, reusing the existing code
path for inrange GEPs.
[clang][bytecode] Use bytecode interpreter in toplevel `Expr::Evaluate*` functions (#226217)
To avoid all the state setup we otherwise do. This also lets us remove
the `EvalInfo::EnableNewConstInterp` flag, which was previously checked
on every `::EvaluateAsRValue` call.
[DirectX] Don't use IRBuilder to create GEPs (#227000)
This switches the DirectX backend to directly call
`GetElementPtrInst::Create()` instead of creating GEPs via IRBuilder.
IRBuilder will soon start canonicalizing GEPs to ptradd representation
(with the first change being
https://github.com/llvm/llvm-project/pull/226425, affecting constant
expressions only), while DirectX needs GEPs to have specific form.
The change implemented here is a minimal one to unblock further work. In
the future, it will become impossible to construct such GEPs through any
API. The DirectX backend needs to switch to using an intrinsic to
represent any non-canonical GEPs it needs. This followup work is tracked
in https://github.com/llvm/llvm-project/issues/227005.
[llvm][llvm-profgen] Fix printing bogus trace info on 32-bit (#226965)
And as a result, fix
llvm/test/tools/llvm-profgen/AArch64/cs-bogus-trace.test when run on Arm
32-bit:
https://lab.llvm.org/buildbot/#/builders/122/builds/59
Fixes #225569 / 28fec94991905d8dd8a08658e9590e13b7fe1af5.
`x` is for an unsigned integer but the argument given for it is a
uint64_t. This worked on 64-bit (maybe it's passed in a register), but
not on 32-bit where we got strange values printed.
Use format_hex instead, so we don't have to choose a format code.
[lldb] Correct types for ScriptedFrame get_function_name get_display_function_name (#226938)
These claimed to return str but the initial value for self.name is None,
and the C++ side returns optional<string>. So I think optional string is
correct here.
[AMDGPU] Fix missed WMMA C-operand co-exec hazard
The gfx1250 WMMA co-execution hazard check treats only A, B and the
SWMMAC index as registers the in-flight MMA still reads. C (src2 of a
non-SWMMAC WMMA) is missing, so a VALU scheduled into the MMA's shadow
can clobber C and the MMA consumes the new value.
This is latent while C is tied to vdst, since the existing D check then
covers it. It miscompiles where the tie does not hold: for
v_wmma_bf16f32_16x16x32_bf16, whose D is narrower than C, and for the
_threeaddr form of any WMMA.
Add src2 to the checked set for non-SWMMAC WMMAs.
[clang] Cache normalized constraints by expression instead of by declaration (#226620)
Sema::NormalizationCache was keyed by the constrained declaration. The
members of a class template specialization are distinct declarations for
every specialization, but they all share the uninstantiated constraint
expressions of the member of the primary template they were instantiated
from. So the constraints of e.g. the constrained constructors of
std::optional, std::span, std::pair or of the members of range adaptors
were normalized from scratch for every single specialization of those
classes that a TU uses. Normalization is expensive: it expands the whole
concept tree below the expression and substitutes the parameter mapping
at every level.
When compiling Chromium, 48%
(chrome/browser/glic/...contents_manager.cc)
to 72% (chrome/browser/ui/views/frame/browser_view.cc) of all
NormalizationCache misses were for expressions that were normalized
before
for a different declaration, and normalization is 4.5% of all compile
[28 lines not shown]
[Polly] Keep proximity dependence exact if simplifying it unbounds its distance (#225840)
Before calling the isl scheduler, `runIslScheduleOptimizer` gists the
dependences
with the iteration domains (`-polly-opt-simplify-deps=yes`, the default
since
a26db470834a, 2012). For a statement that reads a value written in one
iteration
of a loop in all later iterations, the gist also drops the upper bound
of that
loop:
{ S[k, i, k] -> S[k, i, j] : k < j <= n - 1 } becomes { S[k, i, k] ->
S[k, i, j] : j > k }
The proximity distance along j is then unbounded. isl's Pluto-like step
needs a
bound `u·n + w` on every proximity distance, finds no row that satisfies
it, and
[67 lines not shown]
[VectorCombine] Pass flags during IR creation (#193271)
Since commit 777d6b5, VectorCombine has been using InstSimplifyFolder to
simplify vector instructions during IR construction. When creating a new
instruction, InstSimplifyFolder may fold the operation and return an
existing operand instead of emitting a new instruction.
In such cases, copying IR flags to the returned value is incorrect and
may unintentionally propagate flags to pre‑existing instructions,
polluting the original IR. To avoid this, flags should be passed at IR
creation rather than being set after construction. Fix #192607.
[clang] Consistently cache failed constraint normalization (#227086)
Previously, if substituting the parameter mappings of a normalized
constraint failed, Sema::getNormalizedAssociatedConstraints() returned
nullptr, but it stored the partially substituted normal form in
NormalizationCache. So the first lookup for such a declaration failed,
but every later lookup returned the broken normal form, and subsumption
checking and the ambiguous-constraint diagnostics then continued with
it.
I believe this wasn't intentional:
- Before #161671 (e9972debc98c), normalization was a single step, and a
failure was cached as nullptr.
- #161671 added the parameter mapping substitution step. It inserted the
normal form into the cache before substituting, and returned nullptr if
the substitution then failed, leaving the non-null normal form in the
cache.
[29 lines not shown]
[mlir][LLVM] Use a disjoint scope domain when inlining noalias
This matches recent changes to the LLVM inliner.
AI disclosure: Claude wrote the code, I wrote the commit message and
have done initial review.
[mlir][LLVM] Add disjointScopes to AliasScopeDomainAttr
This also updates the MLIR-side inliner to clone disjoint domains
while cloning alias scopes, matching changes to LLVM.
AI disclosure: Claude wrote the code, I wrote the commit message and
looked at the code.
[AMDGPU] Use a disjoint scope domain for merged LDS structs
When lowering LDS values, all the values are mutually disjoint, so we
can use the newly-added disjoint scopes feature to simplify the IR.
AI disclosure: Claude wrote this and I reviewed it and wrote the
commit message
[AMDGPU] Use a disjoint scope domain for noalias kernel arguments
All noalias arguments of a kernel are disjoint with each other, so we
can use a disjoint scope to save on metadata construction.
AI disclosure: Claude wrote this, I looked at it and wrote this
message.