LLVM/project 196fffe — llvm/lib/Target/RISCV RISCVTargetTransformInfo.cpp RISCVSubtarget.cpp, llvm/lib/Target/RISCV/MCTargetDesc RISCVInstPrinter.cpp

[RISCV] Declare command line options in TableGen (#230014)

Move the cl::opts of RISCVCodeGen into RISCVOptions.td, and those of
RISCVDesc and RISCVAsmParser into MCTargetDesc/RISCVMCOptions.td,
registered by LLVMInitializeRISCVTargetMC(). RISCVTargetMachine holds
`const RISCVOptions &CLOpts` and RISCVSubtarget copies it;
RISCVAsmBackend, RISCVInstPrinter, and RISCVTargetStreamer hold `const
RISCVMCOptions &CLOpts`. -riscv-rvv-regalloc (RegisterPassParser) stays
cl::opt.

`llvm-objdump -M emit-x8-as-fp` now sets a file-static flag instead of
writing the cl::opt.

-riscv-br-merging-base-cost now takes effect when given once; it
previously required getNumOccurrences() > 1.

Aided by Opus 5.5
DeltaFile
+132-0llvm/lib/Target/RISCV/RISCVOptions.td
+24-95llvm/lib/Target/RISCV/RISCVTargetMachine.cpp
+15-79llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+17-55llvm/lib/Target/RISCV/RISCVSubtarget.cpp
+15-21llvm/lib/Target/RISCV/MCTargetDesc/RISCVInstPrinter.cpp
+6-29llvm/lib/Target/RISCV/RISCVTargetTransformInfo.cpp
+209-27928 files not shown
+376-41134 files

LLVM/project 3882ac1 — llvm/lib/Target/AMDGPU AMDGPUInstCombineIntrinsic.cpp

Update for comments
DeltaFile
+12-12llvm/lib/Target/AMDGPU/AMDGPUInstCombineIntrinsic.cpp
+12-121 files

LLVM/project 796a67e — llvm/lib/Target/AMDGPU AMDGPUInstCombineIntrinsic.cpp, llvm/test/Transforms/InstCombine/AMDGPU llvm.amdgcn.sudot.ll llvm.amdgcn.dot.ll

[AMDGPU] Fold add of a variable into a zero dot accumulator

When the dot intrinsic has a zero accumulator, no clamp and a single
add user, fold the add operand into the accumulator:
```llvm
  %dot = call i32 @llvm.amdgcn.sdot4(i32 %a, i32 %b, i32 0, i1 false)
  %r = add i32 %dot, %x
=>
  %r = call i32 @llvm.amdgcn.sdot4(i32 %a, i32 %b, i32 %x, i1 false)
```
If %x is defined after the dot in the same block, the dot is moved
down to the add. The fold is skipped across blocks to avoid sinking
the dot into a loop.
DeltaFile
+112-0llvm/test/Transforms/InstCombine/AMDGPU/llvm.amdgcn.dot.ll
+25-11llvm/lib/Target/AMDGPU/AMDGPUInstCombineIntrinsic.cpp
+2-4llvm/test/Transforms/InstCombine/AMDGPU/llvm.amdgcn.sudot.ll
+139-153 files

LLVM/project 34c7f5a — clang/test/Interpreter value-print-temporaries_99994.cpp value-print-temporaries_99998.cpp

[𝘀𝗽𝗿] initial version

Created using spr 1.3.7
DeltaFile
+1-0clang/test/Interpreter/value-print-temporaries_99994.cpp
+1-0clang/test/Interpreter/value-print-temporaries_99998.cpp
+1-0clang/test/Interpreter/value-print-temporaries_99997.cpp
+1-0clang/test/Interpreter/value-print-temporaries_99996.cpp
+1-0clang/test/Interpreter/value-print-temporaries_99995.cpp
+1-0clang/test/Interpreter/value-print-temporaries_99999.cpp
+6-099,994 files not shown
+100,000-0100,000 files

LLVM/project 17f8bd5 — clang/include/clang/Basic BuiltinsAMDGPU.td, clang/test/CodeGenOpenCL builtins-amdgcn-gfx1250-tensor-load-store.cl

[Clang][AMDGPU] Use unsigned int for tensor builtin D# groups (#230714)

The D# tensor descriptor groups of __builtin_amdgcn_tensor_load_to_lds
and __builtin_amdgcn_tensor_store_from_lds are bit fields, not signed
values. D0 was already declared as a vector of unsigned int; D1 through
D4 are now consistent with it.

OpenCL does not allow lax vector conversions, so the tests are updated
to pass unsigned vectors. The generated IR is unchanged.

Reference: https://github.com/llvm/llvm-project/pull/193310

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+13-15clang/test/CodeGenOpenCL/builtins-amdgcn-gfx1250-tensor-load-store.cl
+2-2clang/test/SemaOpenCL/builtins-amdgcn-error-gfx1250-param.cl
+2-2clang/include/clang/Basic/BuiltinsAMDGPU.td
+17-193 files

LLVM/project aed5dbc — lldb/docs conf.py index.md, lldb/docs/man lldb.rst

[lldb][docs] Use project-local documentation links

Replace same-project absolute URLs with relative Markdown links and Sphinx
cross-references so local and archived documentation stays self-contained.
Update generated Python API docstrings accordingly, and enable the absolute
link check for the LLDB docs to keep it that way.
DeltaFile
+4-4lldb/docs/use/lldbdap.md
+4-4lldb/docs/resources/lldbdap-contributing.md
+3-3lldb/docs/man/lldb.rst
+2-2lldb/docs/index.md
+3-0lldb/docs/conf.py
+1-1lldb/include/lldb/API/SBThread.h
+17-147 files not shown
+24-2113 files

LLVM/project dccbcb7 — llvm/docs conf.py

[docs] Enable absolute documentation link checks

Configure the LLVM documentation URL prefixes so the Sphinx build rejects
absolute links to documents in the same project. Clang is already configured.

Part of #214861
DeltaFile
+11-1llvm/docs/conf.py
+11-11 files

LLVM/project 53191d4 — clang/docs ReleaseNotes.md, clang/lib/Parse ParseTemplate.cpp

[clang][Parse] Delay template-id destruction in NTTP default arguments (#230513)

There was a UAF when the lambda appears within NTTP default arguments.

Co-authored-by: Emery Conrad <emery.conrad at chicagotrading.com>
Co-authored-by: Sebastian Schwartz <sebastian.schwartz at chicagotrading.com>
Co-authored-by: Emery Conrad <emery.conrad at chicagotrading.com>
Assisted-by: Claude Code (claude-opus-5-5)
DeltaFile
+26-0clang/test/Parser/cxx2a-constrained-template-param.cpp
+5-0clang/lib/Parse/ParseTemplate.cpp
+4-0clang/docs/ReleaseNotes.md
+35-03 files

LLVM/project 590e335 — utils/bazel/llvm-project-overlay/clang BUILD.bazel

[Bazel] Fixes fa70420 (#230722)

This fixes fa70420c78df1aff63387108f8c357cb0cb8b99c (#229586).

Buildkite error link:
https://buildkite.com/llvm-project/upstream-bazel/builds?commit=fa70420c78df1aff63387108f8c357cb0cb8b99c

Co-authored-by: Google Bazel Bot <google-bazel-bot at google.com>
DeltaFile
+19-1utils/bazel/llvm-project-overlay/clang/BUILD.bazel
+19-11 files

LLVM/project fa70420 — clang/include/clang/CIR/Dialect/IR CIRDialectBytecode.td, clang/lib/CIR/Dialect/IR CIRDialectBytecode.cpp

[CIR] Add bytecode encodings for attributes (#229586)

Without a BytecodeDialectInterface, MLIR encodes a dialect's attributes
through their assembly format, embedding the printed text in the
bytecode. This adds native encodings for 27 CIR attributes (constants,
constant initializers, global views, method/data-member pointers, C++
special-member attributes, and small metadata attributes), modelled on
the LLVM dialect's bytecode support.

On attribute-dense modules this shrinks the bytecode ~30-40% (8.6KB to
5.4KB across the new tests). The room ahead is huge, emitting bytecode
for CIRGenModule.cpp (441MB of CIR text) still exceeds 22GB RSS and 30
minutes with these encodings (64GB and 45 minutes without), since types
keep using the assembly fallback. This is paving towards selfhosting
with bytecode.

Coverage is partial by design: anything not listed keeps using the
assembly-format fallback, exactly as before. Types keep the fallback too
until the type encodings land separately.

    [3 lines not shown]
DeltaFile
+298-0clang/include/clang/CIR/Dialect/IR/CIRDialectBytecode.td
+208-0clang/lib/CIR/Dialect/IR/CIRDialectBytecode.cpp
+73-0clang/test/CIR/Bytecode/global-view.cir
+58-0clang/test/CIR/Bytecode/attributes.cir
+43-0clang/test/CIR/Bytecode/constants.cir
+34-0clang/test/CIR/Bytecode/special-member.cir
+714-07 files not shown
+789-013 files

LLVM/project acd27db — compiler-rt/lib/scudo/standalone quarantine.h, compiler-rt/lib/scudo/standalone/tests quarantine_test.cpp

[scudo] Check quarantine batch bounds in release builds (#230247)

[Attacking Scudo's Quarantine, section 3: Write Where
Ptr](https://un1fuzz.github.io/articles/quarantine_attack.html#a3)
describes corrupting a quarantine batch's `Count` so that enqueue writes
the freed pointer outside the batch. The original [PoC and exploit code
are
here](https://github.com/un1fuzz/scudo_research/tree/main/quarantine_arbitrary_return).
`push_back()` currently guards its index with a debug-only check; an
oversized count also bypasses enqueue's full-batch equality check.

Make that bound a release-build check. Also validate both counts before
deciding whether batches can merge, enforce merge capacity in
production, and check the count before shuffling (which precedes
recycling). The capacity comparison uses subtraction after validating
both operands, avoiding corrupted-count addition wrapping around.

This closes the out-of-bounds enqueue primitive described in section 3.
It does not address section 2's Double Return attack using in-range

    [23 lines not shown]
DeltaFile
+88-0compiler-rt/lib/scudo/standalone/tests/quarantine_test.cpp
+10-4compiler-rt/lib/scudo/standalone/quarantine.h
+98-42 files

LLVM/project 6bcb0a8 — clang/lib/CIR/CodeGen CIRGenBuiltinX86.cpp

[CIR][NFC] Add missing NYI handling for some x86 builtins (#230609)

While doing a recent code review, I noticed that there were some x86
builtins that were incorrectly falling through to code that handles
builtins below them in a switch. This change adds an errorNYI diagnostic
rather than falling through.
DeltaFile
+4-0clang/lib/CIR/CodeGen/CIRGenBuiltinX86.cpp
+4-01 files

LLVM/project edee7da — llvm/test/CodeGen/AMDGPU combine-and-sext-bool.ll or.ll

test(AMDGPU): trim boolean combine coverage
DeltaFile
+0-246llvm/test/CodeGen/AMDGPU/or.ll
+0-30llvm/test/CodeGen/AMDGPU/combine-and-sext-bool.ll
+0-2762 files

LLVM/project 5973b94 — llvm/test/CodeGen/AMDGPU or.ll

test(AMDGPU): drop extra OR test subtargets
DeltaFile
+0-264llvm/test/CodeGen/AMDGPU/or.ll
+0-2641 files

LLVM/project f0c6683 — llvm/test/CodeGen/AMDGPU sra.ll udiv.ll

AMDGPU: Index loads by workitem id in tests shared with r600 (#230667)

These tests relied on -amdgpu-scalarize-global-loads=false to select
vector loads from uniform pointer arguments. They also have r600 run
lines, so keep the kernels and index the input pointers by the workitem
id instead. Also fix shl_v2i16 not using its computed pointers, and
v_shl_32_i64 using the workgroup id. Also add some uniform variants
of some cases.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+5,761-3,731llvm/test/CodeGen/AMDGPU/srem.ll
+6,118-2,908llvm/test/CodeGen/AMDGPU/clmul.ll
+1,327-1,552llvm/test/CodeGen/AMDGPU/mul.ll
+718-729llvm/test/CodeGen/AMDGPU/sdiv.ll
+481-458llvm/test/CodeGen/AMDGPU/udiv.ll
+481-275llvm/test/CodeGen/AMDGPU/sra.ll
+14,886-9,6534 files not shown
+15,809-10,10110 files

LLVM/project 0b73529 — llvm/lib/Target/AMDGPU SIISelLowering.cpp, llvm/test/CodeGen/AMDGPU combine-or-sext-bool.ll or.ll

fix(AMDGPU): guard OR folds with shared conditions

A shared uniform condition still needs materializing after folding a
divergent OR to a select, and sharing can introduce an extra SCC
conversion. The extension's single-use check does not prevent this
code-size regression.

Check the condition's uses and divergence as well. Add regression and
control cases, and consolidate the boolean OR tests in or.ll.
DeltaFile
+633-0llvm/test/CodeGen/AMDGPU/or.ll
+0-286llvm/test/CodeGen/AMDGPU/combine-or-sext-bool.ll
+10-2llvm/lib/Target/AMDGPU/SIISelLowering.cpp
+643-2883 files

LLVM/project 765392a — clang/include/clang/CIR/Dialect/Builder CIRBaseBuilder.h, clang/lib/CIR/Dialect/Transforms CallConvLoweringPass.cpp

[CIR] Pass x87 long double vectors on x86_64 (#230302)

Vectors of x87 long double now go through x86_64 calling-convention
lowering instead of hitting NYI. The ABI library sizes their elements at
128 bits like clang does, so the signatures match classic codegen.

Unions are still moved as a value of their storage type, and a long
double stores only 10 of its 16 bytes. So a union holding an x87 value
next to another member stays NYI unless it's a plain long double and the
other members fit in those 10 bytes. That also stops a silent miscompile
of unions like `union { long double ld; char c[16]; }`.

Assisted-by: Cursor / Claude Opus 5.5
DeltaFile
+346-0clang/test/CIR/CodeGen/call-conv-lowering-x86_64-x87-vector.c
+91-8clang/test/CIR/Transforms/abi-lowering/x86_64-aggregate-nyi.cir
+88-8clang/lib/CIR/Dialect/Transforms/CallConvLoweringPass.cpp
+66-0clang/test/CIR/CodeGen/vector-logical-long-double.cpp
+28-6clang/include/clang/CIR/Dialect/Builder/CIRBaseBuilder.h
+18-0clang/test/CIR/Transforms/abi-lowering/x86_64-vector.cir
+637-221 files not shown
+648-277 files

LLVM/project 5afa2c0 — llvm/include/llvm/CodeGen SDPatternMatch.h, llvm/lib/Target/RISCV RISCVISelLowering.cpp

[RISCV] Reassociate add (add X, (ext Y)), (ext Z) -> add (add (ext Y), (ext Z)), X (#230657)

If Y and Z are extended from the same bitwidth, then we can reassociate
it so they are in the same inner add.

This then allows combineBinOpOfExt to kick in and narrow the inner add
to vwadd.vv. The motivating case is the sad_* kernels in x264.
Reassociating them allows more arithmetic to stay in a smaller LMUL and
takes up to 40% less cycles on sad_16x16 on the K3. This gives a ~3.1%
speedup on 525.x264_r overall.
DeltaFile
+38-0llvm/test/CodeGen/RISCV/rvv/vwadd-sdnode.ll
+35-0llvm/lib/Target/RISCV/RISCVISelLowering.cpp
+22-0llvm/test/CodeGen/RISCV/rvv/fixed-vectors-vwaddu.ll
+5-0llvm/include/llvm/CodeGen/SDPatternMatch.h
+100-04 files

LLVM/project 2e558c7 — llvm/test/CodeGen/AMDGPU sra.ll udiv.ll

AMDGPU: Index loads by workitem id in tests shared with r600

These tests relied on -amdgpu-scalarize-global-loads=false to select
vector loads from uniform pointer arguments. They also have r600 run
lines, so keep the kernels and index the input pointers by the workitem
id instead. Also fix shl_v2i16 not using its computed pointers, and
v_shl_32_i64 using the workgroup id. Also add some uniform variants
of some cases.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+5,761-3,731llvm/test/CodeGen/AMDGPU/srem.ll
+6,118-2,908llvm/test/CodeGen/AMDGPU/clmul.ll
+1,327-1,552llvm/test/CodeGen/AMDGPU/mul.ll
+718-729llvm/test/CodeGen/AMDGPU/sdiv.ll
+481-458llvm/test/CodeGen/AMDGPU/udiv.ll
+481-275llvm/test/CodeGen/AMDGPU/sra.ll
+14,886-9,6534 files not shown
+15,809-10,10110 files

LLVM/project 9436a76 — llvm/lib/Target/AMDGPU SIISelLowering.cpp, llvm/test/CodeGen/AMDGPU combine-or-sext-bool.ll

perf(AMDGPU): fold extended boolean OR to select

Require one-use extensions to avoid replacing an e32 OR with an e64
cndmask while the extended value is still needed elsewhere.
DeltaFile
+286-0llvm/test/CodeGen/AMDGPU/combine-or-sext-bool.ll
+31-12llvm/lib/Target/AMDGPU/SIISelLowering.cpp
+317-122 files

LLVM/project 125bbdb — llvm/lib/Target/AMDGPU GCNHazardRecognizer.cpp, llvm/test/CodeGen/AMDGPU wmma-coexecution-valu-hazards.mir

[AMDGPU] Fix missed WMMA C-operand co-exec hazard (#226503)

The gfx1250 WMMA co-execution hazard check treats only A, B and the
SWMMAC index as registers the in-flight MMA still reads. C (src2 of a
non-SWMMAC WMMA) is missing, so a VALU scheduled into the MMA's shadow
can clobber C and the MMA consumes the new value.

This is latent while C is tied to vdst, since the existing D check then
covers it. It miscompiles where the tie does not hold: for
v_wmma_bf16f32_16x16x32_bf16, whose D is narrower than C, and for the
_threeaddr form of any WMMA.
DeltaFile
+171-2llvm/test/CodeGen/AMDGPU/wmma-coexecution-valu-hazards.mir
+3-4llvm/lib/Target/AMDGPU/GCNHazardRecognizer.cpp
+174-62 files

LLVM/project 500a33c — llvm/lib/Transforms/HipStdPar HipStdPar.cpp, llvm/test/Transforms/HipStdPar math-fixup-libm-decls.ll

HipStdPar: Don't reprocess math library functions as intrinsics (#230603)

All uses were already replaced so this only inserted unused nonsense
declarations by replacing the first 4 characters of the function name
with "__hipstdpar". For example acosh would be replaced with
__hipstdparh.

Co-authored-by: Claude <noreply at anthropic.com>
DeltaFile
+25-0llvm/test/Transforms/HipStdPar/math-fixup-libm-decls.ll
+1-1llvm/lib/Transforms/HipStdPar/HipStdPar.cpp
+26-12 files

LLVM/project 807079e — llvm/lib/Target/AMDGPU SIISelLowering.cpp, llvm/test/CodeGen/AMDGPU setcc-multiple-use.ll combine-and-sext-bool.ll

fix(AMDGPU): fold AND of any-extended booleans

Demanded-bits simplification can turn a boolean sign extension into an
any extension before the target AND combine. Choose sign extension for
the undefined high bits so the AND still folds to a select.

Cover reduced demand, swapped operands, and zero extension. Restore the
fold in setcc-multiple-use after symmetric demanded-bits simplification.
DeltaFile
+46-0llvm/test/CodeGen/AMDGPU/combine-and-sext-bool.ll
+11-7llvm/lib/Target/AMDGPU/SIISelLowering.cpp
+3-4llvm/test/CodeGen/AMDGPU/setcc-multiple-use.ll
+60-113 files

LLVM/project 4de3071 — clang/test/CIR/CodeGen call-conv-lowering-x86_64-x87-vector.c

Test x87 vectors whose element count is not a power of two

Assisted-by: Cursor / Claude Opus 5.5
DeltaFile
+55-0clang/test/CIR/CodeGen/call-conv-lowering-x86_64-x87-vector.c
+55-01 files

LLVM/project 7385e66 — llvm/lib/Support UnicodeNameToCodepointGenerated.cpp, llvm/test/CodeGen/AMDGPU load-global-i16.ll fcanonicalize.ll

Merge branch 'main' into users/adams381/cir-callconv-x86_64-x87-vectors

Assisted-by: Cursor / Claude Opus 5.5
DeltaFile
+32,088-15,710llvm/test/CodeGen/AMDGPU/frem.ll
+23,347-23,371llvm/lib/Support/UnicodeNameToCodepointGenerated.cpp
+5,287-5,656llvm/test/CodeGen/AMDGPU/load-global-i8.ll
+1,502-5,317llvm/test/CodeGen/AMDGPU/fcanonicalize.ll
+3,366-3,401llvm/test/CodeGen/AMDGPU/load-global-i16.ll
+0-4,734llvm/test/tools/llvm-mca/RISCV/tt-ascalon-x/vlseg-vsseg.s
+65,590-58,1892,489 files not shown
+241,791-117,6402,495 files

LLVM/project b0fce46 — lld/MachO ICF.cpp

[Mach-O] Parallelize ICF's section sort. NFC (#230269)

Port of #223216 (ELF) and #229849 (COFF) to Mach-O. Replace the
single-threaded llvm::stable_sort of the ICF inputs with a parallelSort
over packed 64-bit keys, then gather the inputs in key order.

The key is icfEqClass[0] in the high 32 bits and the input index in the
low bits. With --icf=safe_thunks, bit 31 is set for inputs that are not
keepUnique, so keepUnique inputs still come first within each class. The
index breaks remaining ties by original position, so the resulting order
is exactly what stable_sort produced and the output is unchanged.
DeltaFile
+19-10lld/MachO/ICF.cpp
+19-101 files

LLVM/project 7e2d049 — llvm/test/CodeGen/AMDGPU usubo.ll sdwa-peephole.ll

Merge branch 'main' into users/lukel97/dagcombiner-reassociate-adds
DeltaFile
+32,088-15,710llvm/test/CodeGen/AMDGPU/frem.ll
+1,502-5,317llvm/test/CodeGen/AMDGPU/fcanonicalize.ll
+300-1,371llvm/test/CodeGen/AMDGPU/fma-combine.ll
+290-936llvm/test/CodeGen/AMDGPU/llvm.amdgcn.ubfe.ll
+275-670llvm/test/CodeGen/AMDGPU/sdwa-peephole.ll
+291-292llvm/test/CodeGen/AMDGPU/usubo.ll
+34,746-24,296234 files not shown
+41,997-29,139240 files

LLVM/project f6e4a1e — llvm/test/CodeGen/AMDGPU amdgcn.bitcast.96bit.ll amdgcn.bitcast.128bit.ll, llvm/test/CodeGen/X86 atomic-load-store.ll

[SelectionDAG] avoid rounding exact f32->bf16 (STRICT_)FP_ROUND

For f32-to-bf16, DAGTypeLegalizer::SoftPromoteHalfRes_FP_ROUND currently
ignores the Trunc flag and unconditionally lowers ISD::FP_ROUND to
ISD::FP_TO_BF16 as follows:

- On X86 targets with +avx512bf16 or +avxneconvert, ISD::FP_TO_BF16
  lowers to vcvtneps2bf16, which is lossy (it flushes subnormals to
  zero and quietens sNaNs.)

- On targets without hardware bf16 conversion instructions, it performs
  rounding via a runtime libcall (e.g. __truncsfbf2) or a software
  rounding sequence (which can also be lossy, e.g. quieting sNaNs).

The above is unnecessary and lossy; bitcasting to i32 and extracting
the upper 16 bits is cheaper and lossless. Do that.

Tested:


    [11 lines not shown]
DeltaFile
+14,652-16,958llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.1024bit.ll
+4,672-5,733llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.512bit.ll
+1,624-2,114llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.256bit.ll
+893-1,241llvm/test/CodeGen/X86/atomic-load-store.ll
+747-992llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.128bit.ll
+424-549llvm/test/CodeGen/AMDGPU/amdgcn.bitcast.96bit.ll
+23,012-27,58731 files not shown
+26,421-32,77537 files

LLVM/project 737badf — llvm/test/CodeGen/AArch64 andorxor.ll, llvm/test/CodeGen/AMDGPU and.ll or.ll

fix(SelectionDAG): simplify commuted demanded bits

AND/OR demanded-bit simplification uses RHS known bits to simplify the
LHS, but does not retry the RHS using LHS known bits, making
optimizations depend on operand order.

Retry the RHS when the LHS reduces its demanded bits. Add AArch64
and AMDGPU codegen coverage.
DeltaFile
+258-286llvm/test/CodeGen/X86/float-to-arbitrary-fp.ll
+121-135llvm/test/CodeGen/X86/atomic-rm-bit-test.ll
+99-100llvm/test/CodeGen/AMDGPU/float-to-arbitrary-fp-fp8-hw.ll
+78-0llvm/test/CodeGen/AArch64/andorxor.ll
+46-0llvm/test/CodeGen/AMDGPU/or.ll
+44-0llvm/test/CodeGen/AMDGPU/and.ll
+646-52118 files not shown
+756-62524 files

LLVM/project 6c9322d — llvm/lib/Target/X86 X86ISelLowering.cpp, llvm/test/CodeGen/X86 bfloat.ll

[X86] Keep upper-16-bit f32 extractions in XMM for f16/v8i16

Teach combineBitcast and combineVectorInsert to keep this upper-16-bit
f32 extraction in XMM registers rather than bouncing through a GPR.

This paves the way for the following change in SelectionDAG:
  https://github.com/llvm/llvm-project/pull/230557
DeltaFile
+45-0llvm/test/CodeGen/X86/bfloat.ll
+32-0llvm/lib/Target/X86/X86ISelLowering.cpp
+77-02 files