[AMDGPU] Fold fpround of fadd and fsub into v_mad/fma_mixlo and mixhi
MadFmaMixFP32Pats turns (fadd x, y) into (fma x, 1.0, y) and (fsub x, y)
into (fma (-y), 1.0, x) so the mix instructions absorb the operation along
with the f16 or bf16 source modifiers. MadFmaMixFP16Pats and
MadFmaMixFP16Pats_t16 only did this for fmul, so a rounded result still
needed a separate convert for a rounding the mix instructions perform
themselves.
Unlike the f32 patterns these do not require an operand to be an fpextend
of an f16, since an fpround on the result always removes the convert. The
rewrite is exact because the mix instructions round the f32 result again
when they write the 16-bit destination, so it stays f32_to_f16(fma(x, 1.0,
y)).
Assisted-by: Claude Code Opus 5
[AMDGPU] Require flushed FP16 denormals for the mad-mix f16 results
v_mad_mixlo_f16 and v_mad_mixhi_f16 are the unfused gfx900 forms and flush
16-bit denormals, so a denormal half result is written as zero even when the
FP16 mode asks for it to be kept, while the patterns only required the FP32
mode to flush and that is the one a HIP compile turns off on its own.
Assisted-by: Claude Code Opus 5
[NFC][AMDGPU] Add tests for fpround of fadd and fsub feeding the mix instructions (#224909)
The f32 mix patterns already fold fadd and fsub into v_mad_mix_f32 and
v_fma_mix_f32, but the f16 and bf16 forms only fold fmul, so a half or
bfloat result still pays for a separate convert.
Also cover the denormal modes an f16 result depends on and an f16
source feeding an f16 or bf16 mix. The existing mad-mix-lo and
mad-mix-hi functions now flush denormals for every type rather than
for f32 alone.
Assisted-by: Claude Code Opus 5
[LoopInterchange] Avoid overflow in the memory-instruction ratio check (#214920)
populateDependencyMatrix bails out if MaxMemInstrRatio * NumInsts is
less than NumMemInstr * NumMemInstr. Both products are computed in
32-bit unsigned arithmetic and can wrap, e.g., with
-loop-interchange-max-mem-instr-ratio=2147483648 and 20 instructions the
left product wraps to 0 and the pass rejects an otherwise eligible nest.
Compute both products in uint64_t instead, which cannot overflow as all
three operands are 32 bits wide.
Assisted-by: Claude Opus 5, GPT-5.6 Sol, GPT-6 Astra, Claude Fable 5.1.
Co-authored-by: Matt P. Dziubinski <matt-p.dziubinski at hpe.com>
[SandboxVec][LoadStoreVec] Support constant vectors of mixed types
createConstantVector() previously packed the constant store operands
as-is, which only worked when every store had the same element type.
Take the lane type from VecUtils::getCombinedVectorTypeFor() instead and
reinterpret each constant's bits as that type, going through an integer
of matching width via ptrtoint/inttoptr/bitcast. Constants wider than a
lane (e.g. an i64 in an <N x i32>) are split across several lanes in
memory order. Bail out when a constant cannot be reinterpreted, such as
a non-integral pointer or a relocatable address that needs splitting.
Also flatten vector-typed ConstantPointerNull into per-lane nulls, and
bail out on the remaining vector constants such as poison rather than
packing them into the result.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
[lldb] Implement the `__repr__` method for SBStringList. (#224134)
It was annoying to working with `SBStringList` from the lldb python
repl, because printing it out will produce something like.
`<lldb.SBStringList; proxy of <Swig Object of type 'lldb::SBStringList
*' at 0x10556a7b0> >`.
This also implicitly implements `str(SBStringList)`.
Extend the `__getitem__` behaviour to cover integer slices. Add test
cases.
[libc] Move File::close out of line to file.cpp (#224901)
Moved File::close from src/__support/File/file.h into
src/__support/File/file.cpp and removed #include
"src/__support/CPP/new.h" from file.h, matching Dir::close in
src/__support/File/dir.h.
src/__support/CPP/new.h renames global operator delete via __asm__ to
__llvm_libc_delete (which calls free) for the entire translation unit
without renaming operator new. Including new.h in file.h leaked this
replacement operator delete into unit test translation units that
include file.h, causing an AddressSanitizer alloc-dealloc-mismatch in
Test::createCallable during death tests.
Assisted-by: Automated tooling, human reviewed.
[LV] Add tests for early-exit loops with faulting loads (NFC). (#224762)
Pre-committing the tests for using @llvm.speculative.load intrinsic in
LoopVectorize.
Revert "Unlink multicast records when their interface is detached"
This reverts commit e7ea118a97a35e10ded68a7d07d4f39e3e7f57b4.
Opened a new problem instead of the old ones. Need a better fix.
bluhm@ asked for the revert which I take as implicit OK.
Reported-by: syzbot+68c8b44ea4717240232d at syzkaller.appspotmail.com
devel/py-python-box: new port
Box is a Python dictionary subclass that provides dot notation access to
its keys, while behaving like a standard dict in every other respect.
It supports conversion helpers for JSON, YAML, TOML and msgpack, along
with frozen, default and transformation box types.
Co-authored-by: Michael Osipov <michaelo at FreeBSD.org>
PR: 298696
[libc] Fix mlock, mlock2, and munlock signatures (#224905)
Updated libc/include/sys/mman.yaml so the first parameter of mlock,
mlock2, and munlock is const void * rather than void *.
Updated libc/src/sys/mman/mlock2.h and
libc/src/sys/mman/linux/mlock2.cpp so the flags parameter is unsigned
int rather than int.
Assisted-by: Automated tooling, human reviewed.
CodeGen: Move DataLayout computation to CodeGenTargetMachineImpl's ctor
Every target's TargetMachine constructor passed TT.computeDataLayout() as
the DL string argument to the base constructor, duplicating the same call
across all backends. Some backends just didn't bother passing in the ABI
name.
Co-Authored-By: Claude <noreply at anthropic.com> (Claude Opus 4.8)