[InstCombine][NFC] Use uint64_t for ExtractIdx in vector_extract fold (#225156)
## Summary
This changes `ExtractIdx` from `unsigned` to `uint64_t` in the
`Intrinsic::vector_extract` case.
unsigned ExtractIdx = cast<ConstantInt>(Idx)->getZExtValue();
`ExtractIdx` is initialized from the extract index operand via
`getZExtValue()` , which returns `uint64_t`. In the
`get_active_lane_mask` sub-case, it is scaled and compared against the
mask's upper bound:
if (ExtractIdx * ScaleFactor >= ALMUpperBound->getZExtValue())
Here `ALMUpperBound->getZExtValue()` returns `uint64_t`, and
`ScaleFactor` is `unsigned`.
With `ExtractIdx` declared `unsigned`, the value from `getZExtValue()`
[9 lines not shown]
[mlir][openacc] Update remark for privates (#227074)
Update the "thread-private" term to use "Local memory or registers" for
privates remark printing.
Co-authored-by: Yian Su <yians at nvidia.com>
[SelectionDAG] Fix crash scalarizing `FLDEXP` with a legal `<1 x i1>` exponent (#224809)
Fixes #219695
When the result of an `FLDEXP` node is scalarized, we went through the
generic `ScalarizeVecRes_BinOp` handler, which asks for the scalarized
form of both operands. That only works when both operands share the
result's type action. `ldexp` is not a true binop: its exponent has its
own integer vector type, and under AVX-512 `<1 x i1>` is a legal mask
type. It was never scalarized, so it never made it into the table, and
the lookup hit `TableId should be non-zero`. It isn't specific to i1
either: on AArch64, `ldexp <1 x half>, <1 x i32>` hits the same assert
because `v1f16` is scalarized while `v1i32` is widened.
`FLDEXP` now has its own scalarization handler. The FP operand is
scalarized as before, while the exponent is legalized according to its
own type action: its scalarized value is used when one exists, otherwise
the single element is extracted from the vector. This is the same
approach `ScalarizeVecRes_UnaryOp` and `ScalarizeVecRes_SETCC` already
take for operands that don't need scalarizing, and mirrors what
`SplitVecRes_FPOp_MultiType` does for the split case.
[WebAssembly] Add target architecture to object file format (#225979)
Store the target architecture string (`wasm32` or `wasm64`) in a new
`WASM_TARGET_ARCH` subsection in the `linking` custom section and in a
new `WASM_DYLINK_TARGET_ARCH` subsection in the `dylink.0` custom
section.
Previously, architecture detection for object files and shared libraries
relied on heuristics (such as whether memory64 was imported/defined or
presence of 64-bit relocations). For modules that do not access memory,
this resulted in wasm64 objects/shared libraries being incorrectly
detected as wasm32.
With this change:
- `WasmObjectWriter` emits the `WASM_TARGET_ARCH` subsection in
`linking`.
- `wasm-ld` emits `WASM_DYLINK_TARGET_ARCH` when generating shared
libraries and `WASM_TARGET_ARCH` when generating relocatable output
(`-r`).
[6 lines not shown]
[lldb] Speed up evaluating breakpoint conditions by using DIL (#224740)
The goal of this patch is to speed up the evaluation of breakpoint
conditions. Similar to how UserExpression is used to evaluate the
condition expression, DIL lexes and parses the expression once, and then
only evaluates the AST tree on every breakpoint location hit. If DIL
fails at any step, the evaluation falls back to UserExpression, and DIL
doesn't make any new attempts on subsequent breakpoint hits. The
breakpoint default evaluation mode (DIL, UserExpression or DWIM) can be
changed by `target.breakpoints-condition-mode` setting. The mode can
also be changed separately for a specific breakpoint via command line
option `-Z (--condition-mode)` or by SB API
`SBBreakpoint::SetConditionMode`.
[HLSL] Fix -Wunused-variable in #225519 (#227059)
Inline the variable names given the variable names don't add much
clarity and they aren't used anywhere outside of asserts.
X86: Mark the unused defs of FastISel selected div/rem dead (#227053)
The DIV/IDIV instructions define the quotient, the remainder and the
flags, but X86FastISel::X86SelectDivRem only reads one of the quotient
and remainder. Mark the rest dead when building the instruction instead
of relying on later recomputatios.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[Clang] Fix assertion when instantiating a matrix type with an invalid element type (#224331)
Fixes #202744
`Sema::BuildMatrixType` skips the element type check while the element
type is dependent, so `template <typename Y> using matrix_5_5 = Y
__attribute__((matrix_type(5, 5)));` gives a `ConstantMatrixType` whose
element is still unchecked. On instantiation,
`TreeTransform::RebuildConstantMatrixType` called
`ASTContext::getConstantMatrixType` directly, so nothing ever validated
the substituted type. `matrix_5_5<matrix_5_5<float>>` then hit the "need
a valid element type" assertion. The dependent-dimension case
(`matrix_type(R, C)`) was fine because `RebuildDependentSizedMatrixType`
already goes through `BuildMatrixType`.
`RebuildConstantMatrixType` now goes through `Sema::BuildMatrixType` as
well, the same way `RebuildExtVectorType` does for vectors: it takes the
attribute location and wraps the dimensions in integer literals. An
invalid instantiated element type now gets the usual `invalid matrix
element type` error instead of crashing, and the `_BitInt` element check
that this path also skipped is applied too.
[Option] Allow subcommand names as positional arguments
ArgList::getSubCommand() treats every positional argument that matches
a subcommand name as a subcommand, and reports multiple subcommands if
more than one does. That rejects valid command lines where a later
positional argument happens to be spelled like a subcommand, e.g.
git branch clone # create a new branch named "clone"
Add an AllowSubCommandNamesAsPositionals parameter to getSubCommand().
When set, the first positional argument that matches a subcommand name
is the subcommand, and later ones are passed to HandleOtherPositionals
instead of being reported as multiple subcommands.
[ConstantRange] Compute exact no-wrap region w/o materializing CR. (#223969)
Inline logic into makeExactNoWrapRegion() so we do not need to construct
temporary constant ranges.
Together with using makeExecuteNoWrapRegion in SCEV, this improves
compile-time on SCEV-heavy workloads.
Analysis aided by Opus 5.
PR: https://github.com/llvm/llvm-project/pull/223969
[ProfileData] Support merging MD5-based ProfileSymbolList (#226594)
This patch supports merging MD5-based ProfileSymbolList instances and
writing the merged result to an extensible binary profile.
#210235 introduced the MD5-based ProfileSymbolList section in the
Eytzinger layout, where profile merging was initially supported only
from strings to an MD5-based Eytzinger array.
This patch adds DenseSet<uint64_t> GUIDs to ProfileSymbolList to
accumulate 64-bit MD5 hashes when merging MD5-based symbol lists,
while keeping ColdGUIDTable (EytzingerTableSpan) for zero-copy lookups
during compilation. collectGUIDs, contains, and size are updated to
query GUIDs when populated.
RFC:
https://discourse.llvm.org/t/rfc-faster-sample-profile-loading/90957/7
Assisted-by: Antigravity
[clang] Migrate away from PointerUnion::dyn_cast (NFC) (#226653)
Note that PointerUnion::dyn_cast has been soft deprecated in
PointerUnion.h:
// FIXME: Replace the uses of is(), get() and dyn_cast() with
// isa<T>, cast<T> and the llvm::dyn_cast<T>
Literal migration would result in dyn_cast_if_present (see the
definition of PointerUnion::dyn_cast), but this patch uses dyn_cast on
UnsatisfiedConstraintRecord because it is always nonnull.
Specifically, ConstraintSatisfaction::Details only receives nonnull
pointers in the following places:
- ASTNodeImporter::ImportConstraintSatisfaction
- ConstraintSatisfactionChecker::consumeSFINAEFailure
- ConstraintSatisfactionChecker::EvaluateSlow
- ConstraintSatisfactionChecker::Evaluate
- readConstraintSatisfaction
Assisted-by: Antigravity
[llubi] Add support for llvm.speculative.load. (#226839)
Implement the direct form of llvm.speculative.load, where the number of
accessible bytes N is passed as an i64. Only the N accessible bytes are
read from memory and they must be in bounds of the underlying object;
all other bytes are poison. With from_end, the accessible bytes are the
last N bytes of the loaded value. It is UB if N is poison or exceeds the
size of the loaded type.
Support for the oracle form will be added as follow-up.
PR: https://github.com/llvm/llvm-project/pull/226839
[clang][test] Pass -resource-dir in riscv*-toolchain-extra.c tests (#226952)
Fixes issue reported in
https://github.com/llvm/llvm-project/pull/216996#issuecomment-5846218483
These tests check that, with no GCC installation, the driver finds the
linker, crt0 and sysroot relative to its own bin/ directory. The
compiler-rt paths come from the resource directory, which the tests did
not specify, so they relied on the build's CLANG_RESOURCE_DIR placing it
within the fake toolchain tree.
With a relative CLANG_RESOURCE_DIR containing several "..", e.g.
../../../../lib/clang/24 as used by Gentoo, the resource directory
resolves outside of the test tree. This previously passed only because
the unnormalized path still contained "riscv64-nogcc/". Since
99988429d395 removed the dots from the path, the checks fail.
Pass an explicit -resource-dir within the test tree so the tests don't
depend on the build configuration.
Claude Code helped with this.
[clang][CIR] Add missing return stmt in tests (#226534)
This PR makes sure that tests in:
* clang/test/CodeGen/AArch64/sve/
follow the format previously establised in:
* clang/test/CodeGen/AArch64/sve-intrinsics/
[GVN] Preserve vectorization opportunities when PREing loop loads
Loop load PRE can replace an invariant-address load with a loop-carried
PHI and conditional reload after a may-alias store. That scalar recurrence
can prevent vectorization even when runtime alias checks could disambiguate
the original accesses. Subsequent full unrolling then expands scalar code.
Conservatively preserve the header load in innermost loops whose clobber
is a conditional may-alias store through a varying pointer. Keep existing
PRE behavior for invariant-address clobbers, known aliasing, calls, ordered
memory operations, and loops with vectorization disabled or completed.
This is an opportunity heuristic, not a vectorization legality proof.