[AMDGPU][GlobalISel] Keep typed LLTs in RegBankLegalize combines
The S1 cleanup combines in AMDGPURegBankLegalize built new registers
with an untyped s32, missed when lowering switched to extended LLTs.
The regbank combiner later unified these with typed registers, so
constrainRegAttrs retyped e.g. an i32 G_CTPOP result to s32, which
failed instruction selection.
Take the type from the source or destination register instead.
Change-Id: I06cfb7278af045dbaa3ca0cc6c77a3938f4cef01
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[AMDGPU] Validate scale_sel in v_cvt_scale_*
These instructions can be block16 or block32 depending on the target
and scale_sel bits. Block16 is not supported in strict mode.
Re-enable the rest of the instructions in the strict mode but validate
the scale selector.
[AMDGPU][GlobalISel] Fold neg/abs modifiers when mad-mix selects the low half
selectVOP3PMadMixModsImpl re-runs the fneg/fabs match after rewriting Src
to the 32-bit register the 16-bit value is a half of, but only did so on
the isExtractHiElt path, not for isExtractLoElt.
The two halves are not symmetric. An fneg/fabs of a 32-bit float only
touches bit 31, which is the sign bit of the high half, so folding it
into a modifier on the selected high half is correct. The low half's sign
bit is bit 15, which such an fneg/fabs leaves alone, so only a modifier
that acts on each 16-bit element can be folded there. So the source must be
a 2 x 16-bit vector to fold it.
[InstCombine] Handle select-like i1 sext/zexts in canonicalizeClampLike (#227067)
canonicalizeClampLike matches a select of a select. But if the inner
select has an i1 result type it will be canonicalized to a sext or zext
of the condition. In FFmpeg there is a clamping pattern that shifts the
inverted bits instead of the negated, and the inner select ends up being
canonicalized this way: https://godbolt.org/z/5q8bx69x3
This teaches canonicalizeClampLike to use m_SelectLike to catch these
cases too.
[HLSL] Add `InterlockedCompareExchangeFloatBitwise` function and resource methods (#222167)
This PR adds the `InterlockedCompareExchangeFloatBitwise` standalone
function
and resource methods. It completes the interlocked stack.
The operation matches `InterlockedCompareStoreFloatBitwise`, except that
it
reports the previous value. Clang bitcasts both float arguments to `i32`
before the `cmpxchg`, then bitcasts the extracted result back to `float`
before it stores it through the `original_value` reference parameter.
As with the compare-store form, DXC declares this function for 32-bit
float
alone, so the overload set is float only, and the operation reuses the
32-bit
integer DXIL operation, so it works from shader model 6.0 without
capability
bits.
[2 lines not shown]
[HLSL] Add `InterlockedCompareStoreFloatBitwise` function and resource methods (#222166)
This PR adds the `InterlockedCompareStoreFloatBitwise` standalone
function and
resource methods.
`cmpxchg` rejects a float operand, clang therefore
bitcasts both float arguments to `i32` before it emits the `cmpxchg`,
and the
backend only ever sees an integer compare exchange. No DXIL legalization
is
needed for this operation.
DXC declares this function for 32-bit float alone, with no integer
overload,
so the overload set here is float only. `half` and `double` are
rejected.
The operation reuses the 32-bit integer DXIL operation, so it needs no
[6 lines not shown]
[HLSL] Add `InterlockedCompareExchange` function and resource methods (#222165)
This PR adds the `InterlockedCompareExchange` standalone function and
resource
methods. The operation lowers to the same `cmpxchg monotonic` as
`InterlockedCompareStore`.
`InterlockedCompareExchange` reports the previous value, so an
`extractvalue`
reads it from the `cmpxchg` result and stores it through the
`original_value`
reference parameter.
The PR also adds the 64-bit `InterlockedCompareExchange64` methods,
which DXIL
gates on shader model 6.6.
Fixes: https://github.com/llvm/llvm-project/issues/99130
Assisted by: Github Copilot
Merge tag 'mtd/fixes-for-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/mtd/linux
Pull MTD fixes from Miquel Raynal:
"The most important set of fixes are around the handling of the QE bit
in SPI NAND.
There are also a couple of behavioral fixes (mutex issue in SPI-NOR,
spurious bitflips on vf610_nfc, OOB bytes count in SPI NAND and
cfi_cmdset stack usage).
The rest is mostly AI fuzzing results"
* tag 'mtd/fixes-for-7.3-rc6' of git://git.kernel.org/pub/scm/linux/kernel/git/mtd/linux:
mtd: spinand: Do not update the QE bit on devices without one
mtd: spi-nor: core: Fix mutex leak in spi_nor_rww_start_exclusive()
mtd: rawnand: cadence: Initialize IRQ state before requesting IRQ
mtd: rawnand: vf610_nfc: fix false bitflips on reads of erased pages
mtd: rawnand: vf610_nfc: fix reads on chips with more than 64 bytes of OOB
mtd: spinand: fix zero oobavail when no ECC engine is used
[7 lines not shown]
[flang][Evaluate] Fold type-only inquiries of conditional arguments
Fold a reference to an intrinsic inquiry function whose result depends
only on the declared type, kind type parameters, or rank of an argument
when that argument is a conditional argument with a non-constant
condition:
integer, parameter :: k = kind((flag ? a : b))
Every consequent-arg of a conditional argument has the same declared
type and kind type parameters (F2023 C1538) and the same rank unless
all of them are assumed-rank (C1539), so such an inquiry has the same
value whichever consequent-arg is chosen at run time.
FoldOperation(FunctionRef) folds the reference by substituting the
first consequent-arg for the conditional argument in a copy of the
reference and keeps the result when it is a constant.
The qualifying arguments are the dummy arguments that the intrinsic
table marks with ArgFlag::onlyConstantInquiry, which
[34 lines not shown]
[SystemZ][z/OS] Add AMODE to PR symbols
Contrary to the documentation, setting the AMODE at PR symbols is
required. The symptom is that references to variables `optind` and
`optarg` (from include `<getopt.h>`, the LE-provided C runtime)
results in "missing symbol" errors.
Fix is to add AMODE to PrAttr, analog to LdAttr.
TwoAddressInstructions: Fix subranges straddling the INSERT_SUBREG subreg index
Rewriting INSERT_SUBREG into a subregister COPY narrows the def from the
whole register down to one subreg index. The live interval fixup for that
only handled subranges entirely disjoint from the inserted lane mask, so a
partially overlapping subrange keeps a value whose def no longer writes all
of its lanes.
The stale value makes the lanes outside the inserted subregister look
defined by the COPY, so the copy that actually provides them ends up dead
and its value is lost. Refine the subranges against the subreg index before
narrowing the def, leaving each one either fully redefined by the COPY or
untouched by it.
Fixes a miscompile reported on RISC-V and restores the pre-6286f77214be
output of Thumb2/mve-vst2.ll, Thumb2/mve-vst3.ll and PowerPC/dmr-enable.ll.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>