[AIX][libc++] Forward-declare locale functions in aix.h when _XOPEN_SOURCE < 700 (#220508)
On AIX, the following locale function `uselocale` is only declared in
`usr/include/locale.h` when `_XOPEN_SOURCE > 700`.
This causes a failure in the test
`libcxx/test/extensions/posix/xopen_source.gen.py` which tests
`_XOPEN_SOURCE ` against 500, 600 and 700, with the following error:
```
#__locale_dir/support/aix.h:36:72: error: no member named 'uselocale' in the global namespace; did you mean 'setlocale'?
```
The fix here is to forward declare `uselocale` inside `aix.h` and make
it visible.
---------
Co-authored-by: himadhith <himadhith.v at ibm.com>
kk_KZ.PT154 locale: Add missing toupper/tolower mappings.
This was missing a tolower mapping (MAPLOWER) from the lower case
(ASCII) letters into themselves (which is required to work), and worse
was also missing a toupper mapping (MAPUPPER) from the lower case (ASCII)
letters into their upper case equivalents.
Whether the actual (non ASCII) letter upper/lower mappings are all
correct, I am not sure, but I very much doubt it.
Detected and reported by RVP@ in:
https://mail-index.netbsd.org/tech-userlevel/2026/08/15/msg015003.html
www/pomerium: update to 0.33.4
While here:
- Set Pdeathsig on the envoy child process so that it is terminated
when pomerium exits, matching the behavior on Linux.
- Stop enabling reuse_port on FreeBSD: Envoy force-disables it for
TCP listeners on non-Linux platforms and logs a warning for each
listener.
[CodeGen] Report an error for a direct inline asm output in memory
An inline asm output returned by value has no memory to write to, yet a
constraint such as "=rm" picks memory, the most general constraint, as
does "=m". SelectionDAG asserted on that ("Can only indirectify direct
input operands!"), and GlobalISel dereferenced a null pointer. Clang
never emits such an output, since it passes the address of a memory
output, but other IR can. Report "cannot handle direct memory outputs
yet for constraint 'm'" instead, like the other inline asm errors there.
Assisted-by: Claude Opus 5.5
[TargetInstrInfo] Fix folding inline asm operands next to tied operands
foldInlineAsmMemOperand() swapped a register operand for the target's
memory operands with MachineInstr::removeOperand(), which asserts when a
later operand is tied, because moving it would break the tie. Inline asm
lists every input after every output, so folding any operand that comes
before another tied pair asserted, e.g. an "rm" input followed by the
input of a "+r" operand, or one of two "+rm" operands. Without
assertions, the moved operands kept stale tie indices. Untie the
operands, rebuild the operand list, and re-tie the remaining pairs at
their new positions.
It also gave up when the register appears in more than one operand, e.g.
one value passed to two "rm" operands, which left the greedy allocator
unable to spill that value at all. Fold every such operand into the
stack slot.
Finally, take MayLoad from the folded operands rather than from every
read of the register: a folded def whose tied use is another virtual
[8 lines not shown]
[MIR] Serialize the "foldable" inline asm register operand flag
The symbolic MIR syntax for inline asm operand flags dropped
InlineAsm::Flag's RegMayBeFolded bit, which marks a register operand
(from an "rm" constraint) that the register allocator may fold to a
stack slot. Printing MIR and parsing it back therefore silently changed
what the allocator was allowed to do, e.g. with -stop-before and
-run-pass, and a test could only set the bit through a raw numeric flag.
Print the bit as a trailing "foldable", as MachineInstr::print() already
does, and parse it back. A tied use stores its matched operand number in
the same bits, so it never prints the marker.
Assisted-by: Claude Opus 5.5
[CodeGen] Let RegAllocFast lower tied operands by default (#229310)
Following X86 (#228968), the -O0 pipeline no longer runs
TwoAddressInstructionPass.
Opt out:
* AMDGPU: GCNPassConfig::addFastRegAlloc inserts SIWholeQuadMode after
TwoAddressInstructionPass. Without that pass, WWM pseudos reach the
asm printer.
* PowerPC: slightly worse output p9-vinsert-vextract.ll, requiring
RegAllocFast improvement (XXPERM's XA before XTi), the untied read is
allocated first and the tied use needs a copy, adding two vmr.
AArch64 -O0 code has almost no tied operands (94 of 137632 instructions
in sqlite3, mostly INSERT_SUBREG), so the saving is the skipped pass:
https://llvm-compile-time-tracker.com/compare.php?from=680b97545e819edfe94ceb269138061e191685fd&to=8a675992a17edc79aed27959142ca60796a524ce&stat=instructions:u
stage1-aarch64-O0-g instructions:u geomean: -0.43%
Aided by Opus 5.5
[TargetLowering][X86] Prefer 'r' over 'm' for foldable "rm" inline asm operands
An "rm" (register-or-memory) inline asm operand has always resolved to
'm', because getConstraintPreferences() picks the most general
constraint present, and 'm' is more general than 'r'. That's safe, since
memory can't run out, but it forces a value that could stay in a
register through a stack slot even when there's no register pressure
(https://github.com/llvm/llvm-project/issues/20571).
Prefer 'r' instead where the register allocator can fold the register
back to a stack slot when it runs out of registers, and mark the
register operand foldable (InlineAsm::Flag::setRegMayBeFolded()) so it
does. Both allocators can: the greedy allocator folds an operand when it
spills its value, and the fast allocator folds operands up front when
the asm's register operands wouldn't fit.
ParseConstraints() sets AsmOperandInfo::MayFoldRegister for an operand
whose constraint codes are exactly {r, m}, above -O0, on a target that
opts in through the new supportsRegMemInlineAsmFolding() hook, which
[20 lines not shown]
[CodeGen] Report an error for a direct inline asm output in memory
An inline asm output returned by value has no memory to write to, yet a
constraint such as "=rm" picks memory, the most general constraint, as
does "=m". SelectionDAG asserted on that ("Can only indirectify direct
input operands!"), and GlobalISel dereferenced a null pointer. Clang
never emits such an output, since it passes the address of a memory
output, but other IR can. Report "cannot handle direct memory outputs
yet for constraint 'm'" instead, like the other inline asm errors there.
Assisted-by: Claude Opus 5.5
[RegAllocFast] Fold foldable inline asm operands under register pressure
An inline asm register operand marked foldable (from an "rm" constraint)
may be replaced with a stack slot when the register allocator runs out
of registers. The greedy allocator does that when it spills the value.
The fast allocator can't: it assigns the operands of an instruction one
at a time, and folding replaces the instruction. So it reported
"inline assembly requires more registers than available" instead.
Before allocating a block, estimate from each inline asm's own operands
whether they fit in registers, and fold as many foldable registers as
needed to make them fit, so the common case without pressure still gets
a register. Values that only live across the asm don't count, since the
allocator spills them when it needs their registers. The estimate
follows allocateInstruction():
- First all defs get distinct registers, avoiding physreg defs; then all
uses, along with the defs still occupied while the uses are read
(early-clobber and tied defs, see isLiveThroughDef()), get distinct
[25 lines not shown]
[CodeGen] Report an error for a direct inline asm output in memory
An inline asm output returned by value has no memory to write to, yet a
constraint such as "=rm" picks memory, the most general constraint, as
does "=m". SelectionDAG asserted on that ("Can only indirectify direct
input operands!"), and GlobalISel dereferenced a null pointer. Clang
never emits such an output, since it passes the address of a memory
output, but other IR can. Report "cannot handle direct memory outputs
yet for constraint 'm'" instead, like the other inline asm errors there.
Assisted-by: Claude Opus 5.5
[TargetInstrInfo] Fix folding inline asm operands next to tied operands
foldInlineAsmMemOperand() swapped a register operand for the target's
memory operands with MachineInstr::removeOperand(), which asserts when a
later operand is tied, because moving it would break the tie. Inline asm
lists every input after every output, so folding any operand that comes
before another tied pair asserted, e.g. an "rm" input followed by the
input of a "+r" operand, or one of two "+rm" operands. Without
assertions, the moved operands kept stale tie indices. Untie the
operands, rebuild the operand list, and re-tie the remaining pairs at
their new positions.
It also gave up when the register appears in more than one operand, e.g.
one value passed to two "rm" operands, which left the greedy allocator
unable to spill that value at all. Fold every such operand into the
stack slot.
Finally, take MayLoad from the folded operands rather than from every
read of the register: a folded def whose tied use is another virtual
[8 lines not shown]
graphics/embree3: try to fix the build against CMake 4.x
While here, add forgotten type annotation in the similar
commit made to `math/fastops' port earlier.
PR: 299180
Fixes: fa07c9073bc9
nfscl: Fix oddball cases for session slot release
We have identified some cases where silent slot loss can occur
when operations on NFS mounts are aborted. We experience this
when using NFSv4.2, but it likely also occurs with NFSv4.1.
A slot is acquired for compound operations by nfsv4_setsequence()
and freed by newnfs_request(). Any call path that abandons the
compound before reaching newnfs_request() loses the slot permanently.
We identified four call sites where this happens, one of
which where it actually does happen for us in a semi-reproducible
way, which allowed us to develop a candidate patch, attached.
The patch adds one function, nfsv4_freeunsentslot(), to
nfs_clcomsubs.c. It is called from each of the four call
sites: nfsrpc_writerpc(), nfsrpc_writeds(), and two in
nfsrpc_setextattr().
[11 lines not shown]
[AMDGPU] Fix missed WMMA C-operand co-exec hazard
The gfx1250 WMMA co-execution hazard check treats only A, B and the
SWMMAC index as registers the in-flight MMA still reads. C (src2 of a
non-SWMMAC WMMA) is missing, so a VALU scheduled into the MMA's shadow
can clobber C and the MMA consumes the new value.
This is latent while C is tied to vdst, since the existing D check then
covers it. It miscompiles where the tie does not hold: for
v_wmma_bf16f32_16x16x32_bf16, whose D is narrower than C, and for the
_threeaddr form of any WMMA.
[InstCombine] Treat `asin`, `asinh`, `atan` and `cbrt` as odd-functions (#227336)
Add `asin`, `asinh`, `atan` and `cbrt` into the odd-functions list. They
should be treated as odd-functions now.
For #227011
make.conf: Delete the mislabeled ENABLE_SUID_SSH knob
This knob actually controlled the SUID bit of libexec/ssh-keysign, which
is needed for host-based authentication, and it actually requires the
SUID permission to work.
So just remove the ENABLE_SUID_SSH knob and always install ssh-keysign
with SUID. While there, explicitly set BINOWN=root and change BINMODE
to 4555, following both FreeBSD and OpenBSD.
Obtained-from: FreeBSD (commit 0041e47595fad8de5f6c3fd28522e4aa14eef32a)
lib80211: fix build with eXpat 2.9.0
eXpat 2.9.0 deprecates XML_GetCurrentLineNumber() in favour of
XML_GetCurrentLineNumber64(). The new function behaves the same
as the old one but is not prone to 32 bit integer wrap-around.
[MachineOutliner] Attribute outlined calls to the candidate's last call (#229260)
An outlined call stands in for a whole candidate but can carry only one
debug location, so no choice is correct for every instruction it
replaces. Prioritize the location that keeps unwinding as if the code
were not outlined: a return address is symbolized at the preceding
instruction, so unwinding through a call in the candidate resolves the
caller frame at the outlined call. Using the candidate's first location
could attribute that frame to an unrelated inlined callee, as seen in
ASan reports.
Prefer the last call's location. A candidate ending in a call may be
outlined as a thunk whose tail call returns directly past the outlined
call, making the backtrace match the unoutlined code exactly. With
multiple calls, the location can still be exact for only one of them.
Keep the first location when the candidate has no call, its last call
has no line, or the replacement is a tail branch that nothing returns
to.
Follow-up to llvm/llvm-project#224189.