[flang][cuda] Do not use cuf.alloc for derived-type function results (#227863)
Local variables of a derived type with device allocatable components
are allocated in managed memory with cuf.alloc/cuf.free. This also
applied to function results, but the storage of a derived-type function
result is replaced by the caller-provided buffer in the AbstractResult
pass. Allocating it with cuf.alloc is therefore incorrect, and the
matching cuf.free would release memory that is returned to the caller.
Skip function results in needCUDAAlloc so that they are lowered to a
regular fir.alloca. Explicit device/managed/shared/pinned attributes
on the symbol are still honored.
[mlir] Avoid rewriting unreachable blocks in the greedy driver
A rewrite can disconnect a block after the iteration's initial CFG
sweep. Queued operations can develop self-referential SSA uses in
unreachable code, causing crashes or repeated rewrites that prevent the
worklist pass from returning. Track reachability through rewriter
notifications and skip these operations until the next iteration removes
their blocks.
The reproducer in #221152 exposes this gap in the initial sweep added by
#153957 for #153732. #154038 similarly skips unreachable blocks in the
walk-based driver. Reachable graph-region self-cycles in #194824 and
#205064 remain separate; #207185 proposes a fold-specific fix for them.
A forwarding ReachabilityListener sits between the rewriter and the
worklist driver and records, per region, which blocks changed their
terminator and which blocks were removed. Queries happen between
rewrites, once per popped operation, and update a per-region cache of
reachable blocks and their successors incrementally: added edges extend
[41 lines not shown]
[flang][cuda] Keep length parameters when boxing in data transfer conversion (#227883)
When lowering cuf.data_transfer, CUFOpConversion creates descriptors for
non-descriptor operands in emboxSrc, emboxDst and asHLFIREntity. The
shape
was taken from the defining fir.declare/hlfir.declare, but the length
type
parameters were dropped. For a character entity with a non-constant
length
(e.g. an automatic `character(len=l) :: str(n)` assigned to a managed or
device array), this produced a fir.embox of !fir.char<1,?> without
typeparams, which later hit the `!lenParams.empty()` assertion in
EmboxCommonConversion::getCharacterByteSize during FIR to LLVM codegen.
Retrieve the type parameters from the declare alongside the shape and
pass
them when creating the box. Lengths already present in the type are
elided,
matching FirOpBuilder::createBox.
[clang] Avoid stack exhaustion in recursive `constexpr` calls (#201706)
Guard constexpr function-call evaluation with
`runWithSufficientStackSpace` so deeply recursive calls use a fresh
stack before exhausting the current one. This prevents crashes during
both constexpr analysis and constant folding during LLVM IR generation,
including recursive floating-point expressions.
Fixes #201418
Fixes #200673
Assisted by Codex.
[Clang] Initialize bypassed variables w/ trivial-auto-var-init (#181937)
When -ftrivial-auto-var-init=zero or -ftrivial-auto-var-init=pattern is
enabled, variables whose declarations are bypassed by goto or switch
statements were silently left uninitialized. This patch ensures they are
initialized, matching GCC 16's behavior.
The initialization is emitted at the jump source rather than the jump
target. This ensures correctness in loops: a goto whose source and
destination are both inside the variable's scope does not spuriously
reinitialize it, while a goto that actually bypasses the declaration
does. For computed gotos (where jump sources cannot be determined
statically), we fall back to initializing in the entry block.
The simplest example of the old behavior is:
```c
switch (x) {
int y;
case 1:
[15 lines not shown]
[LLVMABI] Classify vectors at their ABI size
The x86-64 classifier compared a vector's payload width against Clang type
sizes, so vectors with padding were misclassified. For example,
`struct { long double __attribute__((vector_size(16))) v; }` coerced to
`<2 x double>` instead of `<1 x x86_fp80>`.
getABISizeInBits() now counts an x87 element at its allocation size, and the
classifier uses it wherever Clang uses getTypeSize() for a vector. This also
applies to vectors with a non-power-of-two element count, bool vectors, and
vectors of sub-byte _BitInt. isIllegalVectorType now sends only __int128
vectors to memory, not _BitInt(128) ones, as Clang does. isSingleElementStruct
is shared, so AMDGPU picks up the fix too.
Assisted-by: Claude Code / Claude Opus 5.5
devel/p5-CPAN-YACSmoke: rework port to avoid Portscout false positives
CPAN-YACSmoke uses an unconventional Major.Minor_Patch versioning scheme,
which can conflict with or break FreeBSD's port version comparison
mechanisms.
Rework the port to use a more standard Semantic Versioning scheme and
report the correct DISTNAME to Portscout, avoiding false positives.
pfctl: print "pass" on nat/rdr/binat rules again
c2d03a920ec7 rewrote the action printing in print_rule() from OpenBSD,
which has no natpass, and dropped the "pass" keyword. A ruleset
loaded from "pfctl -sn" output therefore lost its nat-pass semantics.
Add a parser test covering nat, rdr, rdr log and binat with pass.
Reviewed by: kp
Approved by: kp (mentor)
Fixes: c2d03a920ec7 ("pfctl: fix anchortypes bounds test")
MFC after: 1 week
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D60183