[NVPTX] Properly support acquire/release/acq_rel atomics pre-SM70 by emitting membars (#222449)
acquire/release/acq_rel are not supported semantics on pre-SM70 atomics.
We need to implement these by relaxing the atomic to "relaxed" and then
surounding it with membars. This makes the atomic sequentially
consistent, which is a superset of acquire-release behavior. We already
do this for pre-SM70 atomic that cmpxchg expand, just not for atomics
that otherwise are natively supported.
[mlir][sparse] Register bufferization dialect in stage pass (#227550)
StageSparseOperations may create bufferization.dealloc_tensor while
staging conversions. Register the dialect as a pass dependency and add a
standalone regression test.
[mlir][bufferization] Make equivalent result removal order-independent (#227752)
After bufferization, in-place tensor results commonly become memref
results that are equivalent to function arguments. These results are
redundant and should be removed before later calling-convention
conversions such as buffer-results-to-out-params.
DropEquivalentBufferResults currently visits each function once in
module order. If a caller precedes its callee, rewriting the callee can
expose an equivalent result in the caller after it has already been
visited. This makes the transformation depend on function order and may
require running the pass repeatedly.
Use a caller-driven worklist to revisit affected functions until no more
results can be dropped. Refresh stored call operations after rewriting
them so recursive call graphs can also converge safely. Termination is
guaranteed because each propagation step is triggered by removing a
function result.
[4 lines not shown]
[Github] Automatically derive python version in release-binaries (#227795)
Makes updating the python version easier, especially for renovate.
Assisted By: LLM
[MLIR][Bufferization] Fix IdentityLayoutMap allocation at function boundaries (#227253)
Resolves silent data corruption when passing subviews across function
boundaries under `LayoutMapOption::IdentityLayoutMap`.
### The Problem
When the `IdentityLayoutMap` option is specified for function boundary
bufferization, all function parameters are expected to have a fully
contiguous, zero-offset layout. However, if a caller passes a non-unit
stride or offset view (e.g. the result of `tensor.extract_slice`), the
bufferization pass incorrectly lowered this to a `memref.cast`.
Since `memref.cast` strips layout metadata but leaves the base pointer
unchanged, this caused silent wrong-value loads in the callee (reading
from offset 0 regardless of the actual dynamic offset).
### The Solution
This patch intercepts the operand materialization logic in
`FuncBufferizableOpInterfaceImpl.cpp` (specifically during `CallOp`
[19 lines not shown]
[mlir][ValueBounds] Skip analysis for identical slice components (#226894)
One-Shot Bufferize repeatedly compares subset slices whose offsets,
sizes, and strides often reuse the same SSA values and attributes. Avoid
constructing a ValueBounds constraint set when the two OpFoldResults are
already identical, while preserving the existing solver fallback for
distinct values.
www/linkwarden: New port
Linkwarden is a self-hosted, collaborative bookmark manager that collects,
organizes and preserves webpages. Every saved link is archived as a
screenshot, a PDF, a single file HTML copy and a readable article, so the
content stays available even when the original page is gone.
Co-authored-by: Jochen Neumeister <joneum at FreeBSD.org>
www/linkwarden: New port
Linkwarden is a self-hosted, collaborative bookmark manager that collects,
organizes and preserves webpages. Every saved link is archived as a
screenshot, a PDF, a single file HTML copy and a readable article, so the
content stays available even when the original page is gone.
Co-authored-by: Jochen Neumeister <joneum at FreeBSD.org>
[CIR][SYCL] Enable relocatable device code for SYCL (#226596)
Emit `sycl_external` functions with sycl-module-id, allow -fgpu-rdc
mangling, and embed offload objects in the host.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
pf: take the rules read lock in pf_handle_getrule()
pfctl -sr calls PFNL_CMD_GETRULE once per rule, and
pf_handle_getrule() takes the rules write lock each time, so listing a
ruleset of N rules stops packet processing N times. Only zeroing the
counters (pfctl -z) needs the write lock. Take the read lock
otherwise, as DIOCGETRULENV does.
Reviewed by: kp
Approved by: kp (mentor)
Fixes: 777a4702c591 ("pf: implement addrule via netlink")
MFC after: 1 week
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D60161
pf: take the rules read lock in pf_handle_getrule()
pfctl -sr calls PFNL_CMD_GETRULE once per rule, and
pf_handle_getrule() takes the rules write lock each time, so listing a
ruleset of N rules stops packet processing N times. Only zeroing the
counters (pfctl -z) needs the write lock. Take the read lock
otherwise, as DIOCGETRULENV does.
Reviewed by: kp
Approved by: kp (mentor)
Fixes: 777a4702c591 ("pf: implement addrule via netlink")
MFC after: 1 week
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D60161
pf: leave the epoch to purge unlinked rules
pf_purge_thread() calls pf_purge_unlinked_rules() in the network
epoch, where sleeping is not allowed, and it takes pf_config_lock, an
sx lock. If a rule is being added at the time, the purge thread
can sleep on the lock, which panics with INVARIANTS.
Reviewed by: kp
Approved by: kp (mentor)
Fixes: f92d9b1aad73 ("pflow: import from OpenBSD")
MFC after: 1 week
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D60160
pf: leave the epoch to purge unlinked rules
pf_purge_thread() calls pf_purge_unlinked_rules() in the network
epoch, where sleeping is not allowed, and it takes pf_config_lock, an
sx lock. If a rule is being added at the time, the purge thread
can sleep on the lock, which panics with INVARIANTS.
Reviewed by: kp
Approved by: kp (mentor)
Fixes: f92d9b1aad73 ("pflow: import from OpenBSD")
MFC after: 1 week
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D60160
pf: free an unparsed rule with pf_krule_free() in pf_handle_addrule()
When the PFNL_CMD_ADDRULE message fails to parse, pf_handle_addrule()
frees the rule with pf_free_rule(), which asserts the rules and config
locks (neither is held) and releases references that
pf_ioctl_addrule() has not taken yet. With INVARIANTS this panics on
any parse error; without, a rule address parsed as PF_ADDR_TABLE makes
pfr_detach_table() dereference NULL.
Use pf_krule_free(), as the ioctl paths do.
Reviewed by: kp
Approved by: kp (mentor)
Fixes: e249f5daa41f ("pf: fix memory leak on rule add parse failure")
MFC after: 1 week
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D60104
pf: free an unparsed rule with pf_krule_free() in pf_handle_addrule()
When the PFNL_CMD_ADDRULE message fails to parse, pf_handle_addrule()
frees the rule with pf_free_rule(), which asserts the rules and config
locks (neither is held) and releases references that
pf_ioctl_addrule() has not taken yet. With INVARIANTS this panics on
any parse error; without, a rule address parsed as PF_ADDR_TABLE makes
pfr_detach_table() dereference NULL.
Use pf_krule_free(), as the ioctl paths do.
Reviewed by: kp
Approved by: kp (mentor)
Fixes: e249f5daa41f ("pf: fix memory leak on rule add parse failure")
MFC after: 1 week
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D60104
libpfctl: zero the counters before summing per-chunk results
The chunked table address functions (set, add, del, clr_astats) add
each chunk's result to the caller's counter without initialising it.
pfctl reuses nadd for the number of tables created, so a replace that
also creates the table is off by one:
pfctl -t foo -T replace 192.0.2.1
reports "2 addresses added".
Zero the counters first, as pfctl_test_addrs() already does. Remove
the workaround for the add case from pfctl (da64f6e047b5), which is
no longer needed.
Add a regression test.
Reviewed by: kp
Approved by: kp (mentor)
[4 lines not shown]
libpfctl: zero the counters before summing per-chunk results
The chunked table address functions (set, add, del, clr_astats) add
each chunk's result to the caller's counter without initialising it.
pfctl reuses nadd for the number of tables created, so a replace that
also creates the table is off by one:
pfctl -t foo -T replace 192.0.2.1
reports "2 addresses added".
Zero the counters first, as pfctl_test_addrs() already does. Remove
the workaround for the add case from pfctl (da64f6e047b5), which is
no longer needed.
Add a regression test.
Reviewed by: kp
Approved by: kp (mentor)
[4 lines not shown]
[lldb-dap][test] Let tests run under both stdio and server adapter modes (#227435)
Add create_debug_adapter(), which picks stdio or server mode based on
self.run_as_server. This allows tests to run under both modes when
toggling `LLDBDAP_RUN_AS_SERVER`, rather than being pinned to stdio.
[mlir][linalg] Document and diagnose pack/unpack memref limits (#225773)
Scoped down from the [original
RFC](https://discourse.llvm.org/t/rfc-transformation-support-for-linalg-pack-linalg-unpack-on-memrefs/91832)
per discussion in #225650.
- Replace/add the `// TODO: Support Memref Pack/UnPackOp...` comment
across all sites with a comment stating the actual invariant, pointing
to #225650 for the reasoning.
- Document the invariant in the `Linalg_PackOp`/`Linalg_UnPackOp`
descriptions.
- Emit a dedicated diagnostic from `structured.pack`, `lower_pack`,
`lower_unpack`, and tiling when the target has memref operands, instead
of a generic/silent failure.
- Add test coverage for the new diagnostics, previously untested.
---------
Signed-off-by: Federico Bruzzone <federico.bruzzone.i at gmail.com>
[SlotIndexes] Add queries for stale indexes
An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.
NFC. No caller yet. LiveDebugVariables is next.