LLVM/project 77489c0 — mlir/lib/Dialect/Bufferization/IR BufferizationOps.cpp, mlir/lib/Dialect/Bufferization/Transforms FuncBufferizableOpInterfaceImpl.cpp

[MLIR][Bufferization] Fix IdentityLayoutMap allocation at function boundaries (#227253)

Resolves silent data corruption when passing subviews across function
boundaries under `LayoutMapOption::IdentityLayoutMap`.

### The Problem
When the `IdentityLayoutMap` option is specified for function boundary
bufferization, all function parameters are expected to have a fully
contiguous, zero-offset layout. However, if a caller passes a non-unit
stride or offset view (e.g. the result of `tensor.extract_slice`), the
bufferization pass incorrectly lowered this to a `memref.cast`.

Since `memref.cast` strips layout metadata but leaves the base pointer
unchanged, this caused silent wrong-value loads in the callee (reading
from offset 0 regardless of the actual dynamic offset).

### The Solution
This patch intercepts the operand materialization logic in
`FuncBufferizableOpInterfaceImpl.cpp` (specifically during `CallOp`

    [19 lines not shown]
DeltaFile
+30-0mlir/test/Dialect/Bufferization/Transforms/one-shot-module-bufferize.mlir
+27-0mlir/test/Dialect/Bufferization/canonicalize.mlir
+8-1mlir/lib/Dialect/Bufferization/IR/BufferizationOps.cpp
+3-2mlir/lib/Dialect/Bufferization/Transforms/FuncBufferizableOpInterfaceImpl.cpp
+68-34 files

LLVM/project f1ad1f0 — mlir/lib/Interfaces ValueBoundsOpInterface.cpp, mlir/test/Dialect/Bufferization/Transforms one-shot-module-bufferize-analysis.mlir

[mlir][ValueBounds] Skip analysis for identical slice components (#226894)

One-Shot Bufferize repeatedly compares subset slices whose offsets,
sizes, and strides often reuse the same SSA values and attributes. Avoid
constructing a ValueBounds constraint set when the two OpFoldResults are
already identical, while preserving the existing solver fallback for
distinct values.
DeltaFile
+23-0mlir/test/Dialect/Bufferization/Transforms/one-shot-module-bufferize-analysis.mlir
+4-0mlir/lib/Interfaces/ValueBoundsOpInterface.cpp
+27-02 files

LLVM/project fa45fec — clang/lib/CIR/CodeGen CIRGenModule.cpp, clang/lib/CIR/FrontendAction CIRGenAction.cpp

[CIR][SYCL] Enable relocatable device code for SYCL (#226596)

Emit `sycl_external` functions with sycl-module-id, allow -fgpu-rdc
mangling, and embed offload objects in the host.

Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+74-0clang/test/CIR/CodeGenSYCL/sycl-external.cpp
+74-0clang/test/CIR/CodeGenSYCL/embed-offload-object.cpp
+53-0clang/test/CIR/CodeGenSYCL/gpu-rdc-mangled-name.cpp
+18-0clang/lib/CIR/FrontendAction/CIRGenAction.cpp
+12-4clang/lib/CIR/CodeGen/CIRGenModule.cpp
+231-45 files

LLVM/project 7afdf83 — lldb/packages/Python/lldbsuite/test/tools/lldb_dap testcase.py, lldb/test/API/tools/lldb-dap/breakpoint-assembly TestDAP_breakpointAssembly.py

[lldb-dap][test] Let tests run under both stdio and server adapter modes (#227435)

Add create_debug_adapter(), which picks stdio or server mode based on
self.run_as_server. This allows tests to run under both modes when
toggling `LLDBDAP_RUN_AS_SERVER`, rather than being pinned to stdio.
DeltaFile
+20-24lldb/packages/Python/lldbsuite/test/tools/lldb_dap/testcase.py
+3-4lldb/test/API/tools/lldb-dap/launch/TestDAP_launch_termination.py
+1-1lldb/test/API/tools/lldb-dap/save-core/TestDAP_save_core.py
+1-1lldb/test/API/tools/lldb-dap/launch/TestDAP_launch_no_lldbinit_flag.py
+1-1lldb/test/API/tools/lldb-dap/breakpoint-assembly/TestDAP_breakpointAssembly.py
+26-315 files

LLVM/project 1c7ca86 — mlir/include/mlir/Dialect/Linalg/IR LinalgRelayoutOps.td, mlir/lib/Dialect/Linalg/TransformOps LinalgTransformOps.cpp

[mlir][linalg] Document and diagnose pack/unpack memref limits (#225773)

Scoped down from the [original
RFC](https://discourse.llvm.org/t/rfc-transformation-support-for-linalg-pack-linalg-unpack-on-memrefs/91832)
per discussion in #225650.

- Replace/add the `// TODO: Support Memref Pack/UnPackOp...` comment
across all sites with a comment stating the actual invariant, pointing
to #225650 for the reasoning.
- Document the invariant in the `Linalg_PackOp`/`Linalg_UnPackOp`
descriptions.
- Emit a dedicated diagnostic from `structured.pack`, `lower_pack`,
`lower_unpack`, and tiling when the target has memref operands, instead
of a generic/silent failure.
- Add test coverage for the new diagnostics, previously untested.

---------

Signed-off-by: Federico Bruzzone <federico.bruzzone.i at gmail.com>
DeltaFile
+34-18mlir/include/mlir/Dialect/Linalg/IR/LinalgRelayoutOps.td
+43-0mlir/test/Dialect/Linalg/transform-lower-pack.mlir
+39-1mlir/test/Dialect/Linalg/transform-op-pack.mlir
+27-9mlir/lib/Dialect/Linalg/Transforms/PackAndUnpackPatterns.cpp
+36-0mlir/lib/Dialect/Linalg/TransformOps/LinalgTransformOps.cpp
+24-0mlir/lib/Dialect/Linalg/Transforms/DataLayoutPropagation.cpp
+203-285 files not shown
+258-4411 files

LLVM/project 3b7c3a0 — llvm/include/llvm/CodeGen SlotIndexes.h, llvm/lib/CodeGen SlotIndexes.cpp

[SlotIndexes] Add queries for stale indexes

An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.

NFC. No caller yet. LiveDebugVariables is next.
DeltaFile
+207-0llvm/unittests/CodeGen/SlotIndexesTest.cpp
+29-0llvm/lib/CodeGen/SlotIndexes.cpp
+14-0llvm/include/llvm/CodeGen/SlotIndexes.h
+1-0llvm/unittests/CodeGen/CMakeLists.txt
+251-04 files

LLVM/project f31228b — llvm/include/llvm/CodeGen LiveDebugVariables.h, llvm/lib/CodeGen LiveDebugVariables.cpp

[LiveDebugVariables] Repair stale SlotIndexes

The analysis keeps its indexes from before the first register allocator
until DBG_VALUEs are emitted, by which point passes in between have
erased some of the instructions they point at. Resolve them at the
start of each allocator run and before emitting.

SlotIndexes can then reclaim the entries of erased instructions without
sparing the ones held here, which would have made generated code depend
on -g. Emitted locations are unchanged, except that intervals resolving
to one position now emit a single DBG_VALUE rather than identical
consecutive ones.
DeltaFile
+140-0llvm/lib/CodeGen/LiveDebugVariables.cpp
+63-0llvm/test/DebugInfo/AMDGPU/live-debug-vars-stale-slot-indexes.ll
+8-4llvm/test/DebugInfo/MIR/X86/live-debug-vars-unused-arg-debugonly.mir
+8-0llvm/include/llvm/CodeGen/LiveDebugVariables.h
+5-2llvm/test/CodeGen/X86/debug-spilled-snippet.mir
+5-2llvm/test/CodeGen/X86/debug-spilled-snippet.ll
+229-81 files not shown
+236-87 files

LLVM/project 578ddf9 — clang/lib/Sema SemaARM.cpp, clang/test/Sema builtins-microsoft-arm64.c

[clang][AArch64] Fix range check for _Read/_WriteStatusReg (#227052)

PR #187290 added a range check of `[0x4000, 0x7fff]` to the
`_ReadStatusReg` and `_WriteStatusReg` builtins. Since `op0 == 2`
registers have bit 14 clear and encode below 0x4000, that check was
incorrect.

This change widens the range back to `[0x0, 0x7fff]` and adds a
regression test.

Fixes #226777

Assisted by: Claude Opus 5 (via VS Code).
DeltaFile
+8-2clang/test/Sema/builtins-microsoft-arm64.c
+3-1clang/lib/Sema/SemaARM.cpp
+11-32 files

LLVM/project 595f4e1 — lldb/examples/python formatter_bytecode.py

[lldb] Fix GTE operator typo in formatter_bytecode.py (#224767)
DeltaFile
+1-1lldb/examples/python/formatter_bytecode.py
+1-11 files

LLVM/project 087e855 — llvm/test/CodeGen/AMDGPU loop-header-align-gfx950.mir

AMDGPU: Update test after serializing max-bytes-for-alignment (#227895)

Test added after initial PR
DeltaFile
+2-2llvm/test/CodeGen/AMDGPU/loop-header-align-gfx950.mir
+2-21 files

LLVM/project 2a93d3d — llvm/lib/Target/NVPTX NVPTXSubtarget.h NVPTXInstrInfo.td, llvm/test/Transforms/AtomicExpand/NVPTX atomicrmw-i64-sm20.ll

[NVPTX] Support native 64-bit atomic add/sub pre-SM32 (#222471)

64-bit min/max/and/or/xor require SM32 but add/sub don't.
DeltaFile
+126-0llvm/test/Transforms/AtomicExpand/NVPTX/atomicrmw-i64-sm20.ll
+7-10llvm/lib/Target/NVPTX/NVPTXISelLowering.cpp
+1-2llvm/lib/Target/NVPTX/NVPTXSubtarget.h
+1-2llvm/lib/Target/NVPTX/NVPTXInstrInfo.td
+135-144 files

LLVM/project b4643ba — flang/lib/Lower ConvertVariable.cpp, flang/test/Lower/CUDA cuda-derived.cuf

[flang][cuda] Do not use cuf.alloc for derived-type function results (#227863)

Local variables of a derived type with device allocatable components
are allocated in managed memory with cuf.alloc/cuf.free. This also
applied to function results, but the storage of a derived-type function
result is replaced by the caller-provided buffer in the AbstractResult
pass. Allocating it with cuf.alloc is therefore incorrect, and the
matching cuf.free would release memory that is returned to the caller.

Skip function results in needCUDAAlloc so that they are lowered to a
regular fir.alloca. Explicit device/managed/shared/pinned attributes
on the symbol are still honored.
DeltaFile
+19-0flang/test/Lower/CUDA/cuda-derived.cuf
+4-0flang/lib/Lower/ConvertVariable.cpp
+23-02 files

LLVM/project 62f1dac — mlir/lib/Transforms/Utils GreedyPatternRewriteDriver.cpp, mlir/test/Transforms test-strict-pattern-driver.mlir canonicalize-cfg-reachability.mlir

[mlir] Avoid rewriting unreachable blocks in the greedy driver

A rewrite can disconnect a block after the iteration's initial CFG
sweep. Queued operations can develop self-referential SSA uses in
unreachable code, causing crashes or repeated rewrites that prevent the
worklist pass from returning. Track reachability through rewriter
notifications and skip these operations until the next iteration removes
their blocks.

The reproducer in #221152 exposes this gap in the initial sweep added by
#153957 for #153732. #154038 similarly skips unreachable blocks in the
walk-based driver. Reachable graph-region self-cycles in #194824 and
#205064 remain separate; #207185 proposes a fold-specific fix for them.

A forwarding ReachabilityListener sits between the rewriter and the
worklist driver and records, per region, which blocks changed their
terminator and which blocks were removed. Queries happen between
rewrites, once per popped operation, and update a per-region cache of
reachable blocks and their successors incrementally: added edges extend

    [41 lines not shown]
DeltaFile
+330-13mlir/lib/Transforms/Utils/GreedyPatternRewriteDriver.cpp
+334-0mlir/test/Transforms/greedy-cfg-reachability.mlir
+148-4mlir/test/lib/Dialect/Test/TestPatterns.cpp
+118-0mlir/test/Transforms/canonicalize-dce.mlir
+57-0mlir/test/Transforms/canonicalize-cfg-reachability.mlir
+51-2mlir/test/Transforms/test-strict-pattern-driver.mlir
+1,038-194 files not shown
+1,133-2310 files

LLVM/project 07b8767 — llvm/test/CodeGen/AMDGPU slp-int-to-fp.ll

update slp-int-to-fp.ll
DeltaFile
+324-391llvm/test/CodeGen/AMDGPU/slp-int-to-fp.ll
+324-3911 files

LLVM/project d2566ea — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU narrow-int-to-bfloat.ll narrow-int-to-fp.ll

few fixes
DeltaFile
+90-90llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+24-24llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+10-1llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+124-1153 files

LLVM/project a39b0b2 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp

format
DeltaFile
+1-1llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+1-11 files

LLVM/project a4bd0d8 — llvm/test/CodeGen/AMDGPU slp-int-to-fp.ll

update slp-int-to-fp.ll
DeltaFile
+128-170llvm/test/CodeGen/AMDGPU/slp-int-to-fp.ll
+128-1701 files

LLVM/project eee9dd7 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU narrow-int-to-bfloat.ll narrow-int-to-fp.ll

[AMDGPU] Do not price the extension of a loaded i40, i48 or i56 in int to fp casts

A widened constant or invariant load still pays it.
DeltaFile
+100-100llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+24-24llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+18-6llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+142-1303 files

LLVM/project 10160c8 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Analysis/CostModel/AMDGPU narrow-int-to-bfloat.ll narrow-int-to-fp.ll

[AMDGPU] Price scalar integer to fp casts by source width and sign

Scalar sources between a byte and 31 bits fell to the default cost of one
while the matching vector lanes were already priced, which skewed the
difference SLP weighs a bundle against. Such a source is extended before
the conversion, and what the extension takes depends on the width, on the
sign and on whether the subtarget has SDWA and 16 bit instructions.
Sources narrower than a byte are left alone, because their vector form is
not priced either.
DeltaFile
+388-388llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+72-72llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+29-9llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+489-4693 files

LLVM/project 0d4258a — llvm/test/CodeGen/AMDGPU slp-int-to-fp.ll

[NFC][AMDGPU] Add SLP to asm tests for int to fp casts of loaded odd width integers
DeltaFile
+855-0llvm/test/CodeGen/AMDGPU/slp-int-to-fp.ll
+855-01 files

LLVM/project 440d6f7 — flang/lib/Optimizer/Transforms/CUDA CUFOpConversion.cpp, flang/test/Fir/CUDA cuda-data-transfer-char-len.fir

[flang][cuda] Keep length parameters when boxing in data transfer conversion (#227883)

When lowering cuf.data_transfer, CUFOpConversion creates descriptors for
non-descriptor operands in emboxSrc, emboxDst and asHLFIREntity. The
shape
was taken from the defining fir.declare/hlfir.declare, but the length
type
parameters were dropped. For a character entity with a non-constant
length
(e.g. an automatic `character(len=l) :: str(n)` assigned to a managed or
device array), this produced a fir.embox of !fir.char<1,?> without
typeparams, which later hit the `!lenParams.empty()` assertion in
EmboxCommonConversion::getCharacterByteSize during FIR to LLVM codegen.

Retrieve the type parameters from the declare alongside the shape and
pass
them when creating the box. Lengths already present in the type are
elided,
matching FirOpBuilder::createBox.
DeltaFile
+24-0flang/test/Fir/CUDA/cuda-data-transfer-char-len.fir
+14-3flang/lib/Optimizer/Transforms/CUDA/CUFOpConversion.cpp
+38-32 files

LLVM/project 0a9f172 — clang/docs ReleaseNotes.md, clang/lib/AST ExprConstant.cpp

[clang] Avoid stack exhaustion in recursive `constexpr` calls (#201706)

Guard constexpr function-call evaluation with
`runWithSufficientStackSpace` so deeply recursive calls use a fresh
stack before exhausting the current one. This prevents crashes during
both constexpr analysis and constant folding during LLVM IR generation,
including recursive floating-point expressions.

Fixes #201418
Fixes #200673

Assisted by Codex.
DeltaFile
+31-6clang/lib/AST/ExprConstant.cpp
+16-0clang/test/SemaCXX/constexpr-float-call-stack.cpp
+11-0clang/test/SemaCXX/constexpr-call-stack.cpp
+3-0clang/docs/ReleaseNotes.md
+61-64 files

LLVM/project 982cf1a — clang/docs UsersManual.md, clang/lib/CodeGen VarBypassDetector.h VarBypassDetector.cpp

[Clang] Initialize bypassed variables w/ trivial-auto-var-init (#181937)

When -ftrivial-auto-var-init=zero or -ftrivial-auto-var-init=pattern is
enabled, variables whose declarations are bypassed by goto or switch
statements were silently left uninitialized. This patch ensures they are
initialized, matching GCC 16's behavior.

The initialization is emitted at the jump source rather than the jump
target. This ensures correctness in loops: a goto whose source and
destination are both inside the variable's scope does not spuriously
reinitialize it, while a goto that actually bypasses the declaration
does. For computed gotos (where jump sources cannot be determined
statically), we fall back to initializing in the entry block.

The simplest example of the old behavior is:
```c
switch (x) {
int y;
case 1:

    [15 lines not shown]
DeltaFile
+646-0clang/test/CodeGen/trivial-auto-var-init-bypass.c
+423-24clang/test/CodeGenCXX/trivial-auto-var-init.cpp
+64-8clang/lib/CodeGen/CGDecl.cpp
+48-0clang/docs/UsersManual.md
+35-3clang/lib/CodeGen/VarBypassDetector.cpp
+20-1clang/lib/CodeGen/VarBypassDetector.h
+1,236-364 files not shown
+1,289-3910 files

LLVM/project 4e28757 — clang/test/CodeGen/X86 x86_64-vector-abi-size.c, llvm/lib/ABI/Targets X86.cpp

[LLVMABI] Classify vectors at their ABI size

The x86-64 classifier compared a vector's payload width against Clang type
sizes, so vectors with padding were misclassified. For example,
`struct { long double __attribute__((vector_size(16))) v; }` coerced to
`<2 x double>` instead of `<1 x x86_fp80>`.

getABISizeInBits() now counts an x87 element at its allocation size, and the
classifier uses it wherever Clang uses getTypeSize() for a vector. This also
applies to vectors with a non-power-of-two element count, bool vectors, and
vectors of sub-byte _BitInt. isIllegalVectorType now sends only __int128
vectors to memory, not _BitInt(128) ones, as Clang does. isSingleElementStruct
is shared, so AMDGPU picks up the fix too.

Assisted-by: Claude Code / Claude Opus 5.5
DeltaFile
+253-6llvm/unittests/ABI/X86TargetInfoTest.cpp
+160-0clang/test/CodeGen/X86/x86_64-vector-abi-size.c
+19-23llvm/lib/ABI/Targets/X86.cpp
+26-0llvm/unittests/ABI/AMDGPUTargetInfoTest.cpp
+23-0llvm/unittests/ABI/TypesTest.cpp
+22-0llvm/unittests/ABI/TargetInfoTest.cpp
+503-292 files not shown
+519-338 files

LLVM/project 1c5f27b — llvm/include/llvm/FileCheck FileCheck.h, llvm/lib/FileCheck FileCheck.cpp

[FileCheck] Inline destructors to avoid missing symbols with hidden visibility (#225978)
DeltaFile
+0-4llvm/lib/FileCheck/FileCheck.cpp
+4-0llvm/include/llvm/FileCheck/FileCheck.h
+4-42 files

LLVM/project aafdf4d — llvm/include/llvm/CodeGen LiveDebugVariables.h, llvm/lib/CodeGen LiveDebugVariables.cpp

[LiveDebugVariables] Repair stale SlotIndexes

The analysis keeps its indexes from before the first register allocator
until DBG_VALUEs are emitted, by which point passes in between have
erased some of the instructions they point at. Resolve them at the
start of each allocator run and before emitting.

SlotIndexes can then reclaim the entries of erased instructions without
sparing the ones held here, which would have made generated code depend
on -g. Emitted locations are unchanged, except that intervals resolving
to one position now emit a single DBG_VALUE rather than identical
consecutive ones.
DeltaFile
+140-0llvm/lib/CodeGen/LiveDebugVariables.cpp
+63-0llvm/test/DebugInfo/AMDGPU/live-debug-vars-stale-slot-indexes.ll
+8-4llvm/test/DebugInfo/MIR/X86/live-debug-vars-unused-arg-debugonly.mir
+8-0llvm/include/llvm/CodeGen/LiveDebugVariables.h
+5-2llvm/test/CodeGen/X86/debug-spilled-snippet.mir
+5-2llvm/test/CodeGen/X86/debug-spilled-snippet.ll
+229-81 files not shown
+236-87 files

LLVM/project 001c198 — llvm/include/llvm/CodeGen SlotIndexes.h, llvm/lib/CodeGen SlotIndexes.cpp

[SlotIndexes] Add queries for stale indexes

An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.

NFC. No caller yet. LiveDebugVariables is next.
DeltaFile
+207-0llvm/unittests/CodeGen/SlotIndexesTest.cpp
+29-0llvm/lib/CodeGen/SlotIndexes.cpp
+14-0llvm/include/llvm/CodeGen/SlotIndexes.h
+1-0llvm/unittests/CodeGen/CMakeLists.txt
+251-04 files

LLVM/project 89a879d — flang/lib/Lower/OpenMP Utils.cpp, flang/test/Lower/OpenMP declare-variant-loop-bounds.f90

Fix variant context for lastprivate bounds

Use the owning directive's evaluation when collecting construct ancestors
for loop-control expressions. Lastprivate can re-evaluate bounds after
loop-body lowering leaves a body evaluation current, which otherwise adds
the loop construct before the loop-control context filter runs.
DeltaFile
+43-0flang/test/Lower/OpenMP/declare-variant-loop-bounds.f90
+8-2flang/lib/Lower/OpenMP/Utils.cpp
+51-22 files

LLVM/project 66813ba — llvm/lib/Target/X86 X86LFIRewritePass.cpp, llvm/test/CodeGen/X86 lfi-align-sjlj.ll

X86: Remove redundant SJLJ landing pad alignment in X86LFIRewritePass (#227295)

X86LFIRewritePass separately scanned for blocks holding a call site's landing pad label. 
With SJLJ exception handling the dispatch block reaches those blocks through an indirect 
jump, and they are no longer marked as EH pads by the time this pass runs, so they would 
otherwise be missed.

EmitSjLjDispatchBlock puts those blocks in a jump table, so they are already aligned as jump 
table targets. Dropping it removes the use of the TargetOptions exception model field from 
this in preparation for its removal.

The new test checks the landing pad alignment, which was previously untested.

Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
DeltaFile
+66-0llvm/test/CodeGen/X86/lfi-align-sjlj.ll
+7-18llvm/lib/Target/X86/X86LFIRewritePass.cpp
+73-182 files

LLVM/project 8f0dcb7 — llvm/test/Analysis/CostModel/AMDGPU narrow-int-to-bfloat.ll narrow-int-to-fp.ll, llvm/test/CodeGen/AMDGPU itofp-odd-width-load.ll

[NFC][AMDGPU] Add more tests for int to fp casts of loaded odd width integers

Covers loads of i24, i40, i48 and i56 converted to fp, both the loads
that are split into narrower extending loads and the constant or
invariant ones that are widened to a scalar load.
DeltaFile
+466-0llvm/test/CodeGen/AMDGPU/itofp-odd-width-load.ll
+432-0llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-fp.ll
+130-0llvm/test/Analysis/CostModel/AMDGPU/narrow-int-to-bfloat.ll
+1,028-03 files