[AMDGPU] Extend new hazard CFG walk for wmma instruction support (#229428)
Replaces `getWaitStatesSinceVALU` with `getMaxVALUWindowDeficit` and
aligns `checkWMMACoexecutionHazards` implementation with
`checkMAIHazards90A`. Now elides the visited set CFG walk and uses the
BFS CFG walk instead.
Fixes ROCM-32079
AI Assisted
[OpenMP] Fix debug locations in GPU reductions (#228622)
GPU reduction codegen temporarily switches the IRBuilder to AllocaIP to
create reduction storage. In Clang, AllocaIP points before the unlocated
allocapt marker, which clears the current debug location. Restoring the
insertion point does not restore the debug location, so the following
inlinable runtime calls lack !dbg.
Use InsertPointGuard for temporary alloca insertion regions, preserving
both the caller insertion point and debug location.
Remove redundant saveIP/restoreIP pairs around helper emitters that
already use InsertPointGuard internally to preserve the builder state.
Add an OpenMPIRBuilder unit test that models Clang's allocapt insertion
point and verifies the expected runtime call has a valid debug location.
Assisted by gpt-5.6.
[2 lines not shown]
[clang][Flang][SystemZ] Enable -mbackchain on Flang (#229863)
Enable -mbackchain on Flang now that it supports s390x. This is needed
in order to pass some libomp tests:
libomp :: tasking/omp_untied_taskloop.f90
libomp :: transform/fuse/do-looprange.f90
libomp :: transform/fuse/do.f90
libomp :: transform/tile/do.F90
libomp :: transform/tile/do_2d.f90
libomp :: transform/tile/do_2d_varsizes.f90
libomp :: transform/unroll/heuristic_do.f90
[clang][Sema] Use 64-bit triple in throw-address-space.cpp (#230051)
__ptr32 has no effect on a 32-bit system where addresses are already
32-bit. This means no error and no error means the test fails.
Fixes #224680.
[LLVM][CodeGen][SVE] Improve lowering for v1f32/f64 when NEON is not available. (#229724)
When NEON is not available it is better to scalarise single element
floating-point vectors than widening them to use Streaming-SVE.
Explicitly make v1f64 scalar_to_vector operations always legal, because
we can use scalar instructions and add a combine to avoid "nop" casts.
NAS-144087 / 27.0.0 / Gate SMB block cloning and Veeam shares on SMB_BLOCK_CLONING (by sonicaj) (#19967)
This PR adds changes to replace the SMB_FASTPATH and SMB_VEEAM license
features with a single SMB_BLOCK_CLONING feature, which gates the ZFS
block cloning and integrity streams smb.conf parameters as well as the
VEEAM_REPOSITORY_SHARE purpose. Creating a Veeam repository share
without it now reports the generic entitlement message naming the SMB
block cloning feature.
Companion change: https://github.com/iXsystems/truenas_license/pull/78
renames the enum member. The two have to land in lockstep since the
entitlement modules read `LicenseFeature` members at import time, so
either side fails to start without the other. The truenas_license PR
should merge first with this one right after; CI here will fail until
the middleware CI image picks up the new truenas_license package.
Original PR: https://github.com/truenas/middleware/pull/19885
Co-authored-by: Waqar Ahmed <waqarahmedjoyia at live.com>
[CIR][CodeGen][NFC] Retire the CodeGenUtils.h catch-all header
Moves the last helpers out of `CodeGenUtils.h` into `ClassUtils.h`,
`ModuleUtils.h`, `TargetUtils.h` and `FunctionUtils.h` (`checkTargetFeatures`,
since it came from CodeGenFunction.cpp) and deletes the header. Only moves code
already on main, so it can be dropped on its own.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share requiresAMDGPUProtectedVisibility
Deduplicates `requiresAMDGPUProtectedVisibility` between CIR and classic CodeGen
into `TargetUtils.h`. The shared version takes a bool for "currently hidden" in
place of the `llvm::GlobalValue` and `cir::VisibilityKind` the two callers
passed.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen] Share the EH personality selection logic
Deduplicates the EH personality selection (`getEHPersonality`,
`getCXXEHPersonality`) between CIR and classic CodeGen, taking the classic
implementation. CIR's copy lacked the z/OS, Wasm and GNUstep-on-CygMing cases;
none are reachable in CIR today, so no test changes.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen] Share isStandardLibraryRTTIDescriptor
Deduplicates `isStandardLibraryRTTIDescriptor` between CIR and classic CodeGen
into `ItaniumCXXABIUtils.h`, taking the classic implementation. The two copies
have been equivalent since #227781 filled in the builtin types CIR was missing.
Assisted-by: Claude Code (Claude Fable 5.1).
[CIR][CodeGen][NFC] Share hasExtraNeonArgument (#227260)
Deduplicates `hasExtraNeonArgument` between CIR and classic CodeGen into
`TargetUtils.h`.
Assisted-by: Claude Code (Claude Fable 5.1).
Transforms: Use changeToCall in LowerInvoke
The pass predates changeToCall, which does the same rewrite and is what other
contexts already use. The only difference is the new branch takes the invoke's
DebugLoc.
[SLP]Fix masked gather GEPs cost, skip GEP index seeds (#228918)
Pass the loaded type and the real base sharing to the chain cost of the
masked gather GEPs. Do not use the single-use GEP index chains as seeds:
they were vectorized with all lanes extracted for the scalar addresses.
Part of #227393.
Assisted-by: Cursor
[DAG] Explicitly truncate splat operand in get_active_lane_mask expansion
x86 can't handle the implicit truncation of build_vector operands so the whilewr_8 test case crashes after expanding get_active_lane_mask (stemming from llvm.loop.dependence.war.mask). Fix it by explicitly truncating it.
[lldb][Windows] Report the DebugBreakProcess halt as SIGSTOP (#229770)
`NativeProcessWindows` reports signal 19 `SIGSTOP` on Linux but
`SIGCONT` for Windows targets, so each interrupt lldb sent to issue a
packet prints:
```
Process 26748 stopped and restarted: thread 2 received signal: SIGCONT
```
This makes typing stdin to a debuggee tedious when debugging.
This patch uses
`UnixSignals::CreateForHost()->GetSignalNumberFromName("SIGSTOP")` to
use the proper signal and adds a regression test.
[lldb][Windows] End the debug session before destroying an lldb-server process (#229433)
When the debuggee exits, lldb-server sends the exit to the client and
then destroys its process object, while the thread that reported the
exit is still running code that uses that object. lldb-server then
sometimes crashes on its way out.
The destructor now ends the debug session first: it stops the debugger
thread and waits until that thread is done with the object.
A process that exits before its first stop never started, so its exit is
no longer reported to the server, which never had that process. Launch
failures are already reported as an error, and the extra exit would
otherwise arrive as the reply to the launch request (This occurs in the
TestMissingDll.py test).
In a stress run on Windows, 194 of 320 debug sessions crashed before
this change and none after it.
[BasicAA] Fix miscompilation with setjmp/longjmp due to missing longjmp re-entry paths in alias analysis (#212297)
Fixes [#198967](https://github.com/llvm/llvm-project/issues/198967).
`EarliestEscapeAnalysis::getCapturesBefore` (used by GVN, DSE,
MemCpyOpt) determines whether an object is captured before a given
instruction using `isPotentiallyReachable()/isNotInCycle()`, both of
which only see the forward CFG. In a function containing a
`returns_twice` call (e.g. `setjmp`), a `longjmp` can re-enter at that
call site — a back-edge invisible to both checks.
This causes `BasicAA` to incorrectly return `NoAlias` for a local alloca
whose address was captured on a branch that only runs before the
`longjmp` re-entry. Downstream passes (GVN, DSE) then miscompile the
function by eliminating stores that are still live on the re-entry path.
Fix: `EarliestEscapeAnalysis` now caches whether its function contains a
call that may return twice (`callsReturnsTwiceFn()`, backed by the
existing `Function::callsFunctionThatReturnsTwice()` scan, memoized per
instance). If so, `getCapturesBefore` conservatively treats the object
[23 lines not shown]
jj: update to 0.46.0.
### Release highlights
* Jujutsu can now colocate workspaces besides the default one by creating Git
worktrees. Use `jj workspace add --[no-]colocate` and the setting
`git.colocate` to control this.
### Breaking changes
* The minimum supported `git` command version is now 2.42.0, up from 2.41.0.
`jj workspace add` uses `git worktree add --orphan`, which was added in
2.42.0.
* The minimum supported Rust version (MSRV) is now 1.97.1.
* `jj bisect run` now runs some consistency checks before proceeding to bisect.
This helps ensure that the command can tell good and bad revisions apart,
and that the working copy does go from bad to good over the provided revset.
[117 lines not shown]