[orc-rt] Reconcile check-rt-process-info changes. (#225592)
3648582c0a41 was supposed to make this test more generic, but
accidentally removed it instead.
Reconcile the dropped changes with 3648582c0a41, which added page-size
checking.
[RISCV] Expand scalar FP undef/poison to 0.0 with F extension. (#225567)
This ensures the undef/poison value is always nan-boxed. This is
important if the value ends up being used by a freeze. If the value
isn't properly nan-boxed, it will be treated as a nan in FP contexts
regardless of its lower bits. If the freeze is also cast to an integer,
the lack of nan-boxing will be noticed and the integer will see the real
value of the lower bits.
Fixes #225455.
[orc-rt] Add Windows page size detection (#224520)
Adds Windows page size detection for ORC-RT using GetSystemInfo.
Also updates the process-info regression coverage into a single test
[LLVM][AutoUpgrade] Support default args on overloaded intrinsics (#217859)
This patch extends intrinsic `DefaultValue` auto-upgrade to overloaded
intrinsics.
---------
Signed-off-by: DharuniRAcharya <dharunira at nvidia.com>
[AMDGPU] Port Atomic-Optimizer to use Wave Reduction Intrinsics (#211757)
Currently the atomic optimizer creates reductions via intrinsics,
and introduces new control flows.
Replace this sub-target dependent logic with the existing
wave reduction intrinsics, which get lowered in the backend.
This patch ports the uniform-value and divergent-no-return-value cases.
To port the divergent-with-return-value cases, additional scan intrinsics
will need to be added.
[lld] Add -z mark-plt support for X86_64 (#206002)
Add option [no]mark-plt to enable it. Dynamic linker can change PLT
entries with JMPABS instruction on supported targets.
Ref.:
https://maskray.me/blog/2021-09-19-all-about-procedure-linkage-table#x86-plt-rewriting
Assisted-by: Claude Sonnet 4.6
---------
Co-authored-by: Claude Opus 4.8 <noreply at anthropic.com>
[libc++] Build GoogleBenchmark directly from Lit (#224192)
Previously, we would build GoogleBenchmark against the just-built
library, not against the library being tested. When testing historical
versions of libc++ or other standard libraries, this breaks. So instead
of building Google Benchmark against the just-built library in CMake, do
it from Lit as part of the test suite's configuration.
I'm not a huge fan of using Lit as a poor man's build system and we
should make the CMake test suite self-contained, however this is a step
in the right direction and it removes a major coupling between the test
suite and the regular libc++ build.
[libc++] Pin down the OS version of macOS runners (#225536)
Otherwise, we sometimes end up using runners that don't have Xcode 26.5
on them. Pinning this allows bumping the version explicitly when we are
ready to.
[LV] Narrow truncated inductions in a VPlan transform (#220730)
VPRecipeBuilder::tryToOptimizeInductionTruncate matched a truncate of an
induction phi on the underlying IR, and it built the narrowed
VPWidenIntOrFpInductionRecipe from the recipe behind the truncate's
operand.
VPlanTransforms already answers the same question on VPValues in
getOptimizableIVOf, which returns the header IV whether the operand is
the IV
itself or an add of the IV and its step.
Add VPlanTransforms::narrowInductionTruncates and do the match there,
reusing
getOptimizableIVOf and restricting it to the phi for now. The pass runs
right
after makeCallWideningDecisions, which preserves the ordering against
makeScalarizationDecisions that the recipe builder relied on. Building
the
recipe in the transform also removes the need for the conversion loop to
[3 lines not shown]
[PGO] Add GPU wave counters
Lane counts do not show how often a GPU wave reaches an instrumentation
point. A divergent wave can execute both paths even when few lanes take
one path.
Extend the GPU profiling helper to collect a wave count alongside each
existing lane counter. The first active lane increments the wave counter
once, using the same workgroup sampling decision as the lane counters.
Collect both channels together so shared functions use a consistent
counter layout across translation units. Workgroup sampling controls
profiling overhead.
Carry wave counts through raw and indexed profiles and weighted merging
as a separate channel. Preserve ordinary lane-flow counts and leave
blocks without instrumentation unmeasured.
[libc] Implement dual freestore rotation for baremetal heap (#209811)
Implement dual FreeStore rotation in FreeListHeap under
LIBC_COPT_BAREMETAL_HEAP_ENABLE_FREESTORE_ROTATION option for baremetal
targets.
- Allocations pull from active store; free() quarantines blocks into
non-active store (1 - active).
- On allocation failure in active store, rotate() flips active index and
migrates/coalesces quarantined blocks into the new active store.
- Added 2-bit prev_free tracking in BlockRef metadata to distinguish
freestore indices.
- Added unit smoke tests and updated fuzzer for dual freestore rotation.
Assisted-by: Gemini and Claude based automation tool (human-in-the-loop)
---------
Co-authored-by: Yifan Zhu <yfzhu at google.com>
Co-authored-by: Claude Fable 5.1 <noreply at anthropic.com>
[lldb] Consistently use "null-terminated" across LLDB (NFC) (#224801)
It appears that both spellings are correct, but "null terminator" is far
more common, while "NUL terminator" is technically precise regarding the
ASCII character name. Most common in LLDB was "NULL terminated" which is
the worst of both worlds. This rallies around "null-terminated".
- Adjective -> null-terminated
- Verb -> null-terminate
- Nouns left unhyphenated but lowercased: null terminator, null
termination
[SLP]Fix crash in bool bitmask reduction match after tree vectorization
The known bits of the narrowed reduction leaves were computed after the
tree vectorization, which drops the operands of the vectorized scalars,
so the analysis dereferenced null operands. Match the bitmask form before
the vectorization and pass the result to the emission.
Fixes #225538
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/225568
[mlir][memref] Reject empty collapse_shape and expand_shape reassociation groups (#225348)
`memref.collapse_shape` and `memref.expand_shape` both accept an empty
reassociation group, which is not a valid reassociation: every group
must be a
non-empty, contiguous segment of dimensions.
* `memref.collapse_shape` with an empty group aborts the compiler for a
non-identity source layout. `CollapseShapeOp::verify` forwards the
reassociation to `computeCollapsedLayoutMap`, which calls
`ArrayRef::back()`
on each group and asserts on the empty one:
mlir-opt: llvm/include/llvm/ADT/ArrayRef.h:151:
const T& llvm::ArrayRef<T>::back() const [with T = long int]:
Assertion `!empty()' failed.
With an identity source layout the op is instead silently accepted.
[18 lines not shown]
[RegAlloc] [X86] Enable callee saved register optimization for x86 (#220090)
Enable callee saved register optimization implemented in
RAGreedy::tryAssignCSRFirstTime() for x86. It can replace save/restore
instructions in prologue/epilogue with register spill/reload in cold
blocks or register splits.
Spec cpu 2006 int result with fdo on skylake.
```
regalloc-csr-cost-scale 0 30
400.perlbench 42.0 42.7
401.bzip2 25.5 26.3
403.gcc 42.1 41.4
429.mcf 45.0 44.4
456.hmmer 38.2 38.2
458.sjeng 32.3 32.0
462.libquantum 68.0 68.9
471.omnetpp 26.9 27.3
[5 lines not shown]
[SelectionDAG] Add ISD::ARITH_FENCE to SelectionDAGDumper. (#225485)
We seem to have no consistency on CamelCase or snake_case in node
naming. I've gone with CamelCase to match the nearby nodes, but happy to
change.
[clang] Return early if a value dependent recovery init appeared in constant evaluation context in legacy constant evaluator (#225027)
A recovery default member initializer can be value-dependent even when
the expression referring to the variable is not. Clang should return
early to avoid crash.
This fix the issue found in
https://github.com/llvm/llvm-project/issues/185874#issuecomment-4058045596.
---------
Signed-off-by: yronglin <yronglin777 at gmail.com>
[orc-rt] Simplify bit.h countl_zero and bit_width (#225547)
The old countl_zero algorithm wasn't recognized / optimized by clang on
arm64 or x86-64. Switch to a simpler loop that clang recognizes and
rewrite bit_width in terms of countl_zero. NFCI.
[Driver] Link ubsan_loop_detect with --whole-archive (#225498)
`addSanitizerRuntimes` places sanitizer archives before user object
files on the linker command line.
Because `ubsan_loop_detect` was in `NonWholeStaticRuntimes` without `-u`
symbols, single-pass linkers like GNU `ld.bfd` discarded
`libclang_rt.ubsan_loop_detect.a` before seeing references to
`__ubsan_install_trap_loop_detection` or `__ubsan_is_trap_loop`. Move
`ubsan_loop_detect` to `StaticRuntimes` so it is linked with
`--whole-archive`.
[CIR] Use cleanup active flag with logical operators (#225554)
When temporary expressions are created within a logical binary
operation, we need to use a "cleanup active" flag to guard any cleanups
that are created in the right-hand side of the expression because the
expression may short-circuit and not evaluate the RHS. Failure to do so
had been leading to destructors being called for objects that had never
been constructed.
This fix introduces a regression in destructor call ordering when both
sides of a logical operation create temporaries that require cleanup.
This is a known ordering bug that preceeded this PR but was incidentally
avoided by the previous incorrect handling. The orderig bug will be
fixed in a follow-up change.
Assisted-by: Cursor / various models
[CIR] Fix 'cir.not' lowering behavior for >64 bit size (#225541)
We were only inverting the lower 64 bits because we used the
int64_t/uint64_t overload, which only filled in 64 bits. This patch
replaces that with a 'getAllOnes' of the right size.
[CIR] Lower nobuiltin attribute (#225545)
This causes a problem in tests for global allocation functions, but we
are not currently lowering the 'nobuiltin' attribute to LLVM-IR. This
patch adds the lowering.