[NVPTX] Generalize flexible operands to floating-point instructions (#230258)
Extend the register-or-immediate operand approach to floating-point
instructions and intrinsics.
[CIR][OpenMP] Add support for the OpenMP 'for' directive
This patch adds support for wsloop in ClangIR: the `for` directive and its
combined forms `parallel for` and `target parallel for`. This is lowered to an
omp.wsloop + omp.loop_nest, nested utilizing the existing queue-based
decomposition.
Assisted-by: Cursor / Claude Sonnet 5 High
[SLP]Fix deps for stores ahead of may-throw calls
Make stores-after-may-throw insts control dependent on the next
may-throw instruction.
Fixes #230402
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/230567
[CIR] Fix layout and attributes for non-power-of-two vectors
Struct members, array elements and union storage are now laid out by
alloc size, as in LLVM, so a three-element vector takes its full 16
bytes instead of 12. Pointer differences and the x86_64 va_arg stride
use the alloc size too.
The calling-convention pass now places noundef and nofpclass the way
classic does. A coercion that widens the value drops noundef. Each
half of a flattened argument keeps the argument's attributes. An
indirect argument's pointer no longer carries nofpclass.
Assisted-by: Cursor / Claude Opus 5.5
[llvm][UnicodeData] Check for LIBXML_READER_ENABLED before building UnicodeCharSetsGenerator (#230254)
Since LIBXML_READER_ENABLED is a required feature for building
UnicodeCharSetsGenerator, I added a check in the CMakeLists to ensure
that if there's a static library of libxml2 that library has
LIBXML_READER_ENABLED.
I double checked and this looks like the only required libxml2 feature
for UnicodeCharSetsGenerator.
[libc][math] Optimize exp2f and exp10f hot paths (#224178)
Route common finite inputs directly to the existing degree-5 kernels so
exp2f avoids its exceptional-path frame and exp10f keeps cold handling
out of line. Preserve the proven polynomial arithmetic and special
cases.
Add targeted MPFR regression inputs and official range-based performance
coverage for both functions.
PerfTest medians from alternating baseline/patch runs were:
```
| |baseline | patch | speedup|
+----------------------+----------+-----------+--------+
|exp2f hot normal |2.33805 | 1.92944 | 17.02% |
|exp2f close to one |2.28649 | 1.76252 | 22.78% |
|exp10f hot normal |2.83441 | 2.36588 | 15.80% |
|exp10f close to one |2.89851 | 2.25680 | 20.67% |
[6 lines not shown]
[CIR] Support bool vectors in x86_64 calling-convention lowering (#230280)
Bool vectors are now sized one bit per element, as clang does, and are
passed and returned the way classic CodeGen passes them, including in
records.
Assisted-by: Cursor / Claude Opus 5.5
[MLGO] Do not cap evictions when logging default advisor decisions (#230502)
Otherwise trace collection hits the Regs[CandidatePos].second assertion
on large functions (ones with many evictions)
[mlgo] Allow passing pre-emitc-ed models (#227941)
Support pre-lowering models and then passing them via the exact same mechanism - i.e. `LLVM_MLGO_MODELS`. The extension for the pre-generated ones needs to be `.inc`. High level, this just skips trying to run the mlir toolchain over those. Mixing `.inc` and `.mlir` is supported. The mlir toolchain isn't required unless `.mlir` are passed in the list.
As a result we can test the AOT case in regular builds. We just always append to the `LLVM_MLGO_MODELS`list the test models, with an "ugly" command line flag (a `_test` prefix). Each pass just lists the mlir test model and its corresponding .inc as part of the call to `MLGOLower`.
The bulk of the change is changing tests accordingly, and the addition of the same models we use in the mlir case, but EmitC-ed.
A subsequent change will remove listing the mlir models in llvm-zorg, since they now get auto-appended to the list when the mlir tools are specified. In the interim (after this change lands but before we change zorg) the ml-rel bot won't get red because we register the models under a dfferent name on zorg.
Issue #199007
[mlir][x86] Fix result tracing through loops (#230229)
Fixes the search for the write of a contraction result in the AMX
lowering when the result is passed through loop args.
The value was followed to the wrong loop result, which either found the
wrong write or crashed when the loop has few results.
Assisted-by: Claude
[mlir][ArithToLLVM] Fix index lowering for addui_extended (#223265)
Lowering `arith.addui_extended` with `index` operands fails because the
LLVM
result struct uses the unconverted sum type, even though the operands
have
already been converted.
Convert the sum type before constructing the LLVM result struct. Extend
the
existing index-bitwidth tests to cover `arith.addui_extended` at 32-,
64-, and
128-bit widths, reusing their RUN lines. Rename
`constant-index-bitwidth.mlir`
to `index-bitwidth.mlir` to reflect the broader coverage.
This preserves the operation's existing `index` support and uses the
converted
index width for the overflow intrinsic. Related discussion of `index`
[12 lines not shown]
[lldb][NativePDB] Only list a compile unit's own global variables (#230126)
`SymbolFileNativePDB::ParseVariablesForCompileUnit` adds every global
data symbol of the globals stream to whichever compile unit it's asked
about. The globals stream covers the whole PDB, so every compile unit
reports all globals of the program, including the CRT's, and target
variable list hundreds of them.
This patch only keeps the variables whose owning compile unit is the one
being parsed.
Requires:
- https://github.com/llvm/llvm-project/pull/230132
Fixes `TestTargetVar` with PDB debug info.
rdar://189620344
[lldb][NativePDB] Look up global variables by qualified name (#230134)
`SymbolFileNativePDB::FindGlobalVariables` looks the name up in an index
keyed by basename, so a qualified name such as `A::g_points` never
matches (`target variable A::g_points` fails with `"can't find global
variable"`).
This patch splits the name with
`CPlusPlusLanguage::ExtractContextAndIdentifier`, looks up the basename,
and when the name was qualified only keeps variables whose qualified
name contains it, as `SymbolFileDWARF::FindGlobalVariables` does.
Fixes `TestStaticVariables` with PDB debug info.
rdar://189618458
[lldb][NativePDB] Use the public symbol as the mangled name of globals (#230132)
`CreateGlobalVariable` passes `"::" + name` as the mangled name of every
global, and `Variable::GetName` prefers the mangled name, so every
global shows a leading `:: ((int) ::C::abc = 123, (&::ref = ...)`,
SBValue::GetName() returning "::i")`.
This patch uses the `S_PUB32` at the variable's address as the mangled
name when there is one, and no mangled name otherwise, which matches
what lldb shows for MSVC ABI globals with DWARF. NativePDB shell tests
are updated, and `TestFunctionRefs` now accepts both forms (the DWARF
output depends on the C++ ABI), which also removes its Windows XFAIL.
rdar://189620487
[BOLT][AArch64] Shrink large binary call relaxation test (#230526)
The test intermittently times out on the bolt-aarch64-ubuntu-nfc
builder, so I am making the input binary a tad smaller.
[offload][nfc] Pull OpenMP's InteropTbl out of PluginManager
PluginManager is shared with OpenACC, so the OpenMP interop table moves
to OmpPluginManager in libomptarget. The OpenMP PM is defined there, and
interop cleanup is registered from initRuntime.
[AMDGPU] Form VOPD dot2 pairs with a literal in src1
A V_DOT2 with a register in src0 and a literal in src1 is matched as if
commuted, and commuted when the VOPD pair is built. This recovers the
pairs lost once MachineCSE stopped leaving those literals in src0.
Co-Authored-By: Claude <noreply at anthropic.com>