[lldb] Untangle PlatformDarwinKernel's kext and kernel index lookups (#220318)
PlatformDarwinKernel searches an index of the local filesystem for kexts
and kernels, but each search was interleaved with creating the Module,
updating the Target and falling back to PlatformDarwin, so nothing else
could reuse it. Pull the two searches out so a follow-up can answer
Platform::FindModuleFiles with them.
This change is NFC except GetSharedModuleKernel assigned module_sp
before testing whether the candidate matched and never cleared it, so a
failed search would still return the last *non-matching* module.
[mlir][SCFToControlFlow] Carry LLVM attributes through scf.parallel lowering (#219218)
`ParallelLowering` builds its `scf.for` nest without copying anything from the
`scf.parallel`, so an `llvm.loop_annotation` placed on a parallel loop is
silently dropped before `ForLowering` can move it onto the latch branch.
`scf.for` and `scf.while` already propagate LLVM-dialect attributes via
`propagateLoopAttrs`, so do the same for `scf.parallel`. A multi-dimensional
`scf.parallel` carries a single attribute dictionary but becomes several loops,
so the attributes go to the innermost one, whose latch is where `ForLowering`
attaches the loop metadata.
Tests cover the 1-D case and a 2-D nest, where the outer latch is checked to
stay unannotated. Verified that the new tests fail without the fix and pass with
it, and that the pre-existing expectations in `convert-to-cfg.mlir` are
unchanged.
[BOLT] Fix data race on the shared .dwp DWARF context (#220119)
As noted by labrinea, 775dc9b8bf58 ("[BOLT] Create and release .dwo
DWARF contexts incrementally") releases every DWO context at the end of
readDebugInfo, leaving the bucket threads of the DWARF rewrite to
re-open them on demand. With a .dwp package that moved the first touch
of a shared context into the parallel phase, and multiple threads
compete for it, in a race for the abbrev table, causing intermittent
failures in dwarf5-ftypes-dwp-input-dwo-output.test.
Open the split CUs of a package up front, from a single thread, and
resolve the abbreviation table of every unit in it. This is not relevant
for the non-dwp case, which is unaffected.
[AMDGPU] Fix user SGPR accounting when parsing MIR
Count privateSegmentSize and LDSKernelId as user SGPRs to preserve the kernel
descriptor count across MIR round trips.
Fixes LCOMPILER-2700.
[ADT] Exit doFind on an empty map, not just an unallocated one (#220294)
`doFind` exits early only when no buckets were ever allocated. A map that
had entries and lost them keeps its bucket array. `NumBuckets` is then not
zero, so every lookup hashes the key and probes.
For a `SmallDenseMap` in small mode, `NumBuckets` is the template parameter
`InlineBuckets`, a nonzero constant. The existing check can never fire for
those maps. An empty one hashes and probes on every lookup.
`getNumEntries() == 0` covers both cases. It also subsumes the old check.
There are no entries without buckets, so `Mask = NumBuckets - 1` is still
safe. It is one test either way, so non-empty lookups are unchanged. The
check goes before `getRep()`. Keeping it after costs 0.158% on clang, so
those loads are not sunk past the branch.
| workload | instructions:u |
|---|---:|
| clang compiling 600 LLVM/Clang/MLIR translation units | **-0.028%** |
[8 lines not shown]
[SampleProf] Strict handling suffixes without trailing "." (#220320)
Trailing "." is followed by variable part of the suffix.
As is, for non dot terminated suffix, getCanonicalFnName pass
with "Dit == It" and cuts any suffix with starts with `Suffix`.
E.g. ".cfi" suffix will match will consume "foo.cfi_something" else,
I believe this is unintentional in general and undesired for ".cfi",
which can match ".cfi_jt".
[NFC][DirectX] Fix a memory leak in resource access (#220350)
This code was leaking memory when `HasGetPtr` was false. We would reset
`GetPtrPhi` to `nullptr` but we would never delete the `PHINode` we
created preemptively.
Fix this by consolidating the logic to handle the case where we don't
need a PHI for the getpointer, and simplify by avoiding insertion of the
instruction at all if we aren't going to use it.
Leak reported by ASAN:
https://lab.llvm.org/buildbot/#/builders/24/builds/23635
[nfc][acc] Make two loop utilities non-static and add tests (#220417)
Updates OpenACCUtilsLoop so that two of the internal helpers are
externally callable. Adds unit tests for them.
Revert "[llvm][AArch64] Ensure stack alignment in non-sibcall tail calls with FPDiff" (#220419)
Reverts llvm/llvm-project#217156
I noticed an issue with it right after it landed, and to simplify
cherry-picking, I'm going to revert and re-land. See:
https://github.com/llvm/llvm-project/pull/220406 for the re-land.
[docs] Finish MyST migration for remaining Clang docs (#220381)
This migrates all remaining clang/docs/**rst files to markdown. There
are still 2-3 remaining generated rst files, and I'm working on that
next. The pixel diff shows 0.4045% pixel differences after
longest-common-subsequence vertical alignment, and they all looked
intentional to me. I don't have a good process for serving that HTML, or
I'd share it.
Tracking issue: #201242
See the [migration guide] for more information.
[migration guide]:
https://llvm.org/docs/SphinxQuickstartTemplate.html#markdown-migration-guidelines
This is a stacked PR based on #220380, which will be a standalone commit
that
renames *.rst -> *.md before this PR lands for history preservation
purposes.
This was prepared with rst2myst plus LLM-assisted cleanup.
[docs] Rename remaining Clang docs for MyST migration (#220380)
Tracking issue: #201242
See the [migration guide] for more information.
[migration guide]:
https://llvm.org/docs/SphinxQuickstartTemplate.html#markdown-migration-guidelines
This is the initial straight rename commit. It will probably break the
docs build, but it has to be a separate PR for blame preservation
purposes.
[AMDGPU][SIInsertWaitcnts] Fix soft wait removal with loop-carried deps
The code that checked if a soft waitcnt (such as the one insterted by
a release fence on LDS) was redundant didn't correctly account for the
fact that that, for example, the previous iteration of a loop could
have introduced memory traffic that needs to be waited on. This bug
appeared to be fairly rare in practice (probably due to the
instruction scheduler shuffling around code in t bad form) but it can
happen.
The fix is that, instad of immediately erasing "redundant" waits, we
add them to a set of waits to be erased, and then remove them from the
set if they prove to be truly redundant.
This has the side effect of fixing a correctness issue around the CAS
loops we emit on gfx1250 - the global_inv we emit after the
`s_loadcnt 0x0` is itself a `loadcnt`-able event, and so needs to be
forced to completion before the next iteration of the CAS loop.
[5 lines not shown]
[AMDGPU] Pre-commit tests for loop-carried memory waits (#220356)
SIInsertWaitcnts is currently dropping waits in a
(fence; read; write; branch) loop in some cases. Pre-commit tests to
show the problem.
AI disclosure: test generated by AI, but I poked them into not being
terrible.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply at anthropic.com>