[lldb] Wait for file to start test_attach_commandline_continue_app_exits (#227317)
`test_attach_commandline_continue_app_exits` starts an inferior that
exits 3s after writing its syncfile. Under load the harness and
lldb-server take longer than that to reach `DebugActiveProcess`. On
Windows, the attach fails with `ERROR_ACCESS_DENIED` and lldb-server
exits without connecting, and the test waits 60s in `accept()`.
This patch gives the inferior a `waitfile:` argument and create the file
once the stub has attached.
rdar://188700994
[VPlan] Print metadata via ModuleSlotTracker. (#203982)
Since #220390 (dd7236de), ModuleSlotTracker assigns stable metadata
IDs to metadata not inserted in a module yet.
Reuse a ModuleSlotTracker instance when printing metadata
for recipes. This should guarantee stable, numeric IDs for metadata in all
cases, and we should no longer need the `vplan-print-metadata` option.
PR: https://github.com/llvm/llvm-project/pull/203982
[lldb][Windows] Improve process_is_running check (#227308)
The current implementation of `process_is_running` uses
[`tasklist`](https://learn.microsoft.com/en-us/windows-server/administration/windows-commands/tasklist)
on Windows. It prints each CSV field through ShowMessage, which calls
`GetConsoleOutputCP`. From a benchmark at desk, when under load, this
call takes `~85ms` instead of `~10us`. This causes tests to time out
when listing a large amount of processes when running check-lldb.
This patch asks the kernel about the one process instead of listing all
of them with tasklist.
rdar://188700994
[flang][codegen] Report a shape or slice cg-rewrite cannot read
cg-rewrite folds a fir.shape, fir.shape_shift, fir.shift or fir.slice
into the code-gen form by reading it through its defining op. A value
that has none cannot be folded. For a slice this went unreported: the
rewrite dropped it and produced a descriptor for the whole array rather
than the section it names. For a shape it reached a cast on a null
defining op.
Report it instead, and say which operand. Rebuilding the value covers a
block argument, but not every case: a slice chosen by an arith.select
has no single value to take apart.
[flang][codegen] Rebuild compile-time-only block arguments in cg-rewrite (#227624)
A fir.shape, fir.shape_shift, fir.shift or fir.slice only describes an
array at compile time. cg-rewrite reads one through its defining op and
folds it into the code-gen form, so none of these types has an LLVM
lowering and none is expected to reach codegen.
A pass that merges two blocks differing only in such a value passes it
as a block argument instead, and a block argument has no defining op.
The fold then finds nothing to read: a slice is dropped, leaving a
descriptor for the whole array rather than the section, and a shape
reaches a cast on a null defining op.
Rebuild the value in the block. The operands behind it are integers,
which can be block arguments, so take those as arguments, forward them
along each branch, and build the value from them at the top of the
block. The merge is kept and no block is duplicated.
graphics/lensfun: unbreak the port's build against CMake 4.x
While here, merge USES+=pathfix into already existing patch
to reduce the number of post-edit backups files.
PR: 298953
[SCEV] Remove expensive inverted reasoning from isImpliedCond. (#227302)
The constructing the inverted expressions and the additional reasoning
is quite expensive (0.35% for CTMark O3), for marginal gain (3 small
regressions on
https://github.com/dtcxzyw/llvm-opt-benchmark-nightly/pull/1494)
I think that should allow us to spend compile-time on SCEV on areas with
higher impact.
2 of those could be recovered by constant-based reasoning about the
inverted condition.
Compile-time improvements:
stage1-O3: -0.35%
stage1-ReleaseThinLTO: -0.32%
stage1-ReleaseLTO-g: -0.28%
stage1-aarch64-O3: -0.31%
stage2-O3: -0.32%
[5 lines not shown]
[ValueTracking] Support multiple predecessors in willNotFreeBetween (#223580)
Previously, `willNotFreeBetween` only walked backward along a linear
chain of single-predecessor blocks, bailing out whenever control flow
branched and merged (eg., conditional `if-else` blocks before a loop
preheader).
This PR generalises the backward walk to use a worklist over all
predecessors from `CtxBB` to `AssumeBB` (bounded by
`MaxInstrsToCheckForFree`). This enables LICM( and possibly other
passes) to prove dereferenceability and hoist invariant loads across
merging control flow when no path contains a freeing instruction.
---------
Co-authored-by: Florian Hahn <flo at fhahn.com>
Co-authored-by: Antonio Frighetto <me at antoniofrighetto.com>
[lldb][Windows] Improve process_is_running check (#227308)
The current implementation of `process_is_running` uses
[`tasklist`](https://learn.microsoft.com/en-us/windows-server/administration/windows-commands/tasklist)
on Windows. It prints each CSV field through ShowMessage, which calls
`GetConsoleOutputCP`. From a benchmark at desk, when under load, this
call takes `~85ms` instead of `~10us`. This causes tests to time out
when listing a large amount of processes when running check-lldb.
This patch asks the kernel about the one process instead of listing all
of them with tasklist.
rdar://188700994
[AMDGPU] Fix missed WMMA C-operand co-exec hazard
The gfx1250 WMMA co-execution hazard check treats only A, B and the
SWMMAC index as registers the in-flight MMA still reads. C (src2 of a
non-SWMMAC WMMA) is missing, so a VALU scheduled into the MMA's shadow
can clobber C and the MMA consumes the new value.
This is latent while C is tied to vdst, since the existing D check then
covers it. It miscompiles where the tie does not hold: for
v_wmma_bf16f32_16x16x32_bf16, whose D is narrower than C, and for the
_threeaddr form of any WMMA.
Add src2 to the checked set for non-SWMMAC WMMAs.
epoch: Fix use-after-free in epoch_trace_report()
epoch_trace_report() assigned the return value of RB_INSERT() back to
the new element. When two threads report the same stack concurrently,
the loser's RB_INSERT() returns the element already in the tree, and
that element was freed while still linked, leaking the new allocation.
The next lookup touches freed memory; KASAN catches it as a
use-after-free.
Keep the return value separate and free the new element instead. The
thread that won the race prints the report, so return without printing
it a second time.
Reviewed by: markj
Fixes: 173c062a569b ("Improve EPOCH_TRACE")
MFC after: 1 week
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D60162
[SCEV] Use howManyLessThans to implement howManyGreaterThans. (#226846)
Update howManyLessThans to support inverting the analyzed comparison and
analyze LHS > RHS as ~LHS < ~RHS.
howManyLessThans should now handle all cases howManyGreaterThans did
(and a few more, see improvements in
https://github.com/dtcxzyw/llvm-opt-benchmark-nightly/pull/1443).
howManyLessThans now accepts an Inverted argument, which analyzes ~LHS <
~RHS, without materializing the inverted operands explicitly, which can
pessimize results.
I tried to update all code paths where this is possible without too many
changes. A few code paths need bigger changes, and are skipped for now.
PR: https://github.com/llvm/llvm-project/pull/226846
[Offloading] Add support for compressed OffloadBinary types (#222774)
Summary:
Offload binaries are used to store many heterogenous architectures into
a singel offloading blob. These lists can get very large so this PR adds
the option to compress them with the LLVM provided compression
libraries.
The implementation is quite simple, we simply compress all the buffers
after the header into a single compressed blob, then re-construct the
header. Extracting is the reverse.
The biggest change is that the offload binary now **owns** the memory,
whereas before we simply took a reference to it. This is necessary
because the decompression must create new memory compared to what the
user provided. This adds an extra copy internally, but it also
simplifies the V2 additions.
This does not wire up any clang/HIP support, just providing the
[6 lines not shown]
net/frr: Report failed configuration reloads (#5748)
* net/frr: Report failed configuration reloads
Preserve the frr-reload exit status so the service controller can report reload failures. Import the frr-reload.log into the standard FRR log at debug level and expose that log directly in the Routing menu.
* Add changelog
[AMDGPU] Stop trying to link AMDGPU ASan from compiler-rt (#227388)
Summary:
As I am building sanitizers in the standard `compiler-rt` location, we
still need to maintain compatibility with the old DeviceRTL usage.
Ideally this will be replaced in the near term, but for now we should
suppress the compiler from trying to link a `-lclang_rt.asan` that
doesn't exist.
---------
Co-authored-by: Matt Arsenault <arsenm2 at gmail.com>
[lldb][test] Change vCont handling in TestEmptyThreadList.py (#227703)
The `def vCont` handler is triggered only for `vCont;c`, not other
`vCont...`. This test only sometimes sends this, so sometimes it
produced a UnexpectedPacketException.
(though it "passes" still, somehow)
To fix this, declare a vCont handler for the colon c version, and let
`def other` handle the rest.
[AMDGPU] Price scalar integer to fp casts by source width and sign
Scalar sources between a byte and 31 bits fell to the default cost of one
while the matching vector lanes were already priced, which skewed the
difference SLP weighs a bundle against. Such a source is extended before
the conversion, and what the extension takes depends on the width, on the
sign and on whether the subtarget has SDWA and 16 bit instructions.
Sources narrower than a byte are left alone, because their vector form is
not priced either.
[AMDGPU] Price narrow integer to bfloat vector casts
A vector lane of 9 to 15 or 17 to 31 bits converted to bfloat got the
generic cost, which leaves out the rounding. Such a lane is converted to
f32 first like any other narrow lane, so price it as the f32 conversion of
the same source plus the rounding. Lanes of 8 and 16 bits keep their cost.
[AMDGPU] Model the cost of the expanded integer to/from floating point casts
No instruction converts to or from a 64 bit integer, and narrow vector
lanes are converted one at a time. Price these expansions by the FP64
rate, sdwa and 16 bit instruction support, and price i33 to i63 like
i64 and bf16 like f32 plus rounding.
Assisted-by: Claude Code Opus 5