[orc-rt] Rename CommandLineParser to OptionParser. (#220461)
"OptionParser" better describes what the class does: it parses a list of
option strings, which need not come from a command line (e.g. options
read from a config file, or constructed in a unit test).
[RISCV] Fix instruction combine of vwaddu.wv and vabd[u].vv (#220143)
The existing combination of `vwaddu.wv` and `vabd.vv/vabdu.vv` may
trigger a stack dump: The source operands of `vwaddu.wv` are not
commutative. Swapping these operands can generate an invalid `VWABDA_VL`
node and cause ISel to panic. This PR fixes the issue and only allows
the instruction with same `mask` and `vl` to be combined.
[ValueTracking][RISCV] Derive range and pow2 for vlenb CSR reads (#219924)
__riscv_vlenb() lowers to a read of the RISC-V vlenb CSR via
llvm.read_register("vlenb"), which returns VLEN/8. Teach
ValueTracking to recognize that read on a RISC-V target and reason
about its result:
* getRangeForIntrinsic() reports the value range of VLENB. RVV
requires VLEN to be a power of two in [32, 65536] (Zvl32b is
the smallest vector extension), so VLENB is in [4, 8192]
regardless of any attribute. This stays sound for Zvl32b,
whose VLEN (32) is not representable as an integer vscale
(VLEN / RVVBitsPerBlock). When the function carries a
vscale_range attribute the subtarget's VLEN is a multiple of
RVVBitsPerBlock (64), so VLENB = vscale * RVVBytesPerBlock
tightens the range.
* isKnownToBeAPowerOfTwo() reports that VLENB is a power of two.
This lets InstCombine fold vlenb comparisons such as
[11 lines not shown]
[ADT][SelectionDAG] Unique SDNodes in a UniquingSet (#220195)
CSEMap keys SDNodes by a serialized FoldingSetNodeID. Every lookup
profiles the opcode, the VT list and the operands into a buffer, and
every probe rebuilds the candidate's whole profile to compare against
it.
UniquingSet compared a key against Info::getKey(*N), building a key from
the candidate on every probe. Add isEqual to the Info contract,
defaulting to that comparison, and expose a FoldingSetNodeID's
accumulated words as a FoldingSetNodeIDRef so an Info that carries one
can hash it.
Key nodes by SDNodeKey, which holds those fields directly and leaves
only what AddNodeIDCustom contributes serialized in a tail, which is
empty for most opcodes. A lookup site keeps its ID.AddXxx lines: the key
forwards them to the tail. isEqual matches the key against a candidate's
fields and builds only the tail; an assertions build cross-checks the
result against the profile both would have had.
Aided by Opus 5
Revert "[IR] `replaceUsesWith` when aliasing a GV and the use is `dso_local_equivalent`" (#220450)
Reverts llvm/llvm-project#220336
Breaks tests after #220339 and supposedly there are still other
outstanding issues.
[AMDGPU] Fix user SGPR accounting when parsing MIR
Count privateSegmentSize and LDSKernelId as user SGPRs to preserve the kernel
descriptor count across MIR round trips.
Fixes LCOMPILER-2700.
[orc-rt] Don't parse program name in CommandLineParser::parse (#220454)
The iterator-based CommandLineParser::parse method expected a program
name as the first argument in the list, but did not do anything with its
value (it was simply skipped). Update it to expect argument strings only
(no program name as first element) and document the new expectation.
The (argc, argv) overload of parse is renamed to parseAsMainArgs. It
still expects a program name as the first argument and will return an
Error if it is not present.
This allows CommandLineParser::parse to be used in contexts where there
is no natural program name (e.g. in unit tests, and options read from
config files).
[orc-rt] Rename test/unit/utils -> test/unit/tools. (#220408)
This aligns the unit test directory with the corresponding header
directory (orc-rt-internal/tools).
[C++20] [Modules] merge the type for anony enum from import and #include (#214121)
Close https://github.com/llvm/llvm-project/issues/213299
Ideally, we shall merge the new enum with the old enum when we creating
the new enum. But the enum is anonymous and the typedef's name come
after the enum body, it is too late to merge them. This is the choice 10
years ago: a523022b5384d7a0901beea7a5f36ee9c09ba339. Actually what we're
merging is the typedef decls.
Then https://github.com/llvm/llvm-project/pull/114240 removes the logic
to remove the new ED. This the direct trigger for the above issue of
ambiguous look ups.
We choose to fix the problem by setting the type of new enum to the type
of the old enum to fix the ambiguous lookup issue.
[flang][runtime] Use CONVERT='SWAP' in endian.f90 test (#220185)
Follow-up to #218302.
The test currently uses CONVERT='BIG_ENDIAN'.
On big-endian systems such as PPC/AIX, this does not actually
exercise endian conversion.
Switch the test to CONVERT='SWAP' so that byte swapping is
performed regardless of the host byte order.
[lld][MachO] Support Objective-C class stubs
Teach Mach-O objc stubs to synthesize class-message stubs that load the class object, selector, and objc_msgSend target.
lld already supports the Apple clang _objc_msgSend$<selector> stub form, even though upstream clang does not currently expose a driver or cc1 flag for emitting it. Apple clang also emits _objc_msgSendClass$<selector>$_OBJC_CLASS_$_<class>; handling that form completes the existing selector-stub support.
Cover local classes, dylib classes, archives, dynamic lookup via -U, missing class symbols, malformed names, unsupported architectures, and dead stripping.
[lldb] Untangle PlatformDarwinKernel's kext and kernel index lookups (#220318)
PlatformDarwinKernel searches an index of the local filesystem for kexts
and kernels, but each search was interleaved with creating the Module,
updating the Target and falling back to PlatformDarwin, so nothing else
could reuse it. Pull the two searches out so a follow-up can answer
Platform::FindModuleFiles with them.
This change is NFC except GetSharedModuleKernel assigned module_sp
before testing whether the candidate matched and never cleared it, so a
failed search would still return the last *non-matching* module.
[mlir][SCFToControlFlow] Carry LLVM attributes through scf.parallel lowering (#219218)
`ParallelLowering` builds its `scf.for` nest without copying anything from the
`scf.parallel`, so an `llvm.loop_annotation` placed on a parallel loop is
silently dropped before `ForLowering` can move it onto the latch branch.
`scf.for` and `scf.while` already propagate LLVM-dialect attributes via
`propagateLoopAttrs`, so do the same for `scf.parallel`. A multi-dimensional
`scf.parallel` carries a single attribute dictionary but becomes several loops,
so the attributes go to the innermost one, whose latch is where `ForLowering`
attaches the loop metadata.
Tests cover the 1-D case and a 2-D nest, where the outer latch is checked to
stay unannotated. Verified that the new tests fail without the fix and pass with
it, and that the pre-existing expectations in `convert-to-cfg.mlir` are
unchanged.
[BOLT] Fix data race on the shared .dwp DWARF context (#220119)
As noted by labrinea, 775dc9b8bf58 ("[BOLT] Create and release .dwo
DWARF contexts incrementally") releases every DWO context at the end of
readDebugInfo, leaving the bucket threads of the DWARF rewrite to
re-open them on demand. With a .dwp package that moved the first touch
of a shared context into the parallel phase, and multiple threads
compete for it, in a race for the abbrev table, causing intermittent
failures in dwarf5-ftypes-dwp-input-dwo-output.test.
Open the split CUs of a package up front, from a single thread, and
resolve the abbreviation table of every unit in it. This is not relevant
for the non-dwp case, which is unaffected.
[AMDGPU] Fix user SGPR accounting when parsing MIR
Count privateSegmentSize and LDSKernelId as user SGPRs to preserve the kernel
descriptor count across MIR round trips.
Fixes LCOMPILER-2700.
[ADT] Exit doFind on an empty map, not just an unallocated one (#220294)
`doFind` exits early only when no buckets were ever allocated. A map that
had entries and lost them keeps its bucket array. `NumBuckets` is then not
zero, so every lookup hashes the key and probes.
For a `SmallDenseMap` in small mode, `NumBuckets` is the template parameter
`InlineBuckets`, a nonzero constant. The existing check can never fire for
those maps. An empty one hashes and probes on every lookup.
`getNumEntries() == 0` covers both cases. It also subsumes the old check.
There are no entries without buckets, so `Mask = NumBuckets - 1` is still
safe. It is one test either way, so non-empty lookups are unchanged. The
check goes before `getRep()`. Keeping it after costs 0.158% on clang, so
those loads are not sunk past the branch.
| workload | instructions:u |
|---|---:|
| clang compiling 600 LLVM/Clang/MLIR translation units | **-0.028%** |
[8 lines not shown]