[AMDGPU] Promote wmma_f32_16x16x128_f8f6f4 to scaled version on gfx1250-strict (#230253)
Added scale factor is 0 which does not change the result. "0" is a
special value that maps to the exponent "1.0" when passed as a constant.
[CIR] Lower non-power-of-two vectors in CallConvLowering (#230087)
Summary:
- The x86_64 classifier now sizes a vector at its ABI size (#227889), so
the bridge no longer needs to reject vectors whose width is not a power
of two.
- The SSEUP walk rounds vector widths the same way, so a union holding
such a vector is still refused at sizes with no coerce type.
- This issue came up while enabling `CallConvLowering` for AMDGPU in
#220197: the gfx950/gfx1250 `transpose-load` builtins return
three-element vectors, and
`CIR/CodeGenHIP/builtins-amdgcn-gfx950-read-tr.hip` and
`builtins-amdgcn-gfx1250-load-tr.hip` fail on the gate.
Related to issue: #220471
Assisted by : claude opus 5.5
[Clang] Mark scoped_atomics with !noalias.addrspace(private)
The HIP specification marks atomics on thread private memory as UB.
Scoped atomics used within a HIP context are also considered UB,
unless explicitley specified via a command line argument.
These are now annotated with !noalias.addrspace(5) for amdgpus,
to avoid an expensive runtime check.
[NFC][RegisterPressure] Add RegisterOperands::restoreLivenessFlags helper (#229952)
Factor the
clear-stale-read-undef-flags-then-recompute-from-LiveIntervals pattern,
currently duplicated in `GCNIterativeScheduler::restoreLivenessFlags`
and `GCNSchedStrategy::modifyRegionSchedule`, into a shared
`RegisterOperands::restoreLivenessFlags` helper, and convert both AMDGPU
call sites to it with no behavior change. The helper also accepts an
optional register filter for targeted recomputation.
Requested as a prerequisite refactor in review of #227897, which will
stack on top of this and switch its new helper to the shared one.
[CIR] Implement BranchOpInterface for switch.flat (#230340)
cir.switch.flat is the CFG-form terminator produced by CIR flattening,
but unlike cir.br and cir.brcond, it doesn't implmenet BranchOpIterface.
As a result, generic MLIR control-flow infrastructure could not tell
which operands are forwarded to which successor's block arguments,. For
example, SCCP pass. When the switch condition is a know constant, SCCP
can know that which sucessor block is actually taken. Before, SCCP
didn't understand switch.flat. It treated every sucessor as reachable,
which forbid it to fold anything.
Implmenet the interface the same way llvm.switch do now.
Assisted-by: Claude # Tests
[flang][cuda] Copy type descriptors for array descriptors in device code (#230305)
The CUFDeviceGlobal pass copies the type descriptors of derived types
used
in device procedures into the GPU module. For fir.embox, it only looked
at
the memref type after stripping the reference, so a scalar derived type
was
handled but an array of derived type (e.g. passing the section `a(1:1)`
of
an assumed-size array to an assumed-shape dummy) was not. The type
descriptor was then missing from the GPU module and FIR-to-LLVM codegen
failed with "runtime derived type info descriptor was not generated".
Use the element type of the resulting box to find the derived type,
which
covers both scalars and arrays. Handle fir.rebox the same way, since
reboxing a non-polymorphic derived type also needs the type descriptor
in
codegen.
[RISCV][GlobalISel] Promote f16 to f32 for G_FPTOSI/G_FPTOUI without Zfh (#230025)
Use the fcvt.s.h conversion with Zfhmin and the __extendhfsf2 libcall
without it, then convert from f32, matching SelectionDAG. This also
fixes failures to legalize fptoui/fptosi from half to i32 without Zfh.
[CIR] Implement BranchOpInterface for switch.flat
Switch.flat should minic llvm.switch as they are both CFG
formed. Implement the same mechanism as what we have in llvm.switch
and add ArrayRef<Attribute> override when we actually know the operands.
Assisted-by: Claude # Tests
[Clang][CodeGen] Fix clang codegen eh cleanup (NFC) (#226743)
NFC for the following reasons:
- `DK_none` and `DK_cxx_destructor` can never reach
`pushLifetimeExtendedDestroy` at either of the two call sites. `DK_none`
is filtered out by the `if (...)` check. `DK_cxx_destructor ` only
occurs for a C++ class, but the code runs only if
`!getLangOpts().CPlusPlus`.
- For the other three (`DK_objc_strong_lifetime`,
`DK_objc_weak_lifetime`, and `DK_nontrivial_c_struct`), the boolean flag
controls whether the remaining not-yet-destroyed elements of an array
get destroyed in `emitArrayDestroy` when the destroyer call itself
throws when destroying an array element. But the flag is irrelevant here
because the destroyer functions for the three never throw: they are all
called via `EmitNounwindRuntimeCall`, which produces a plain `call`
rather than an `invoke`.
[ASan][Darwin] Support gapless shadow layout for iOS 27.0
When the shadow can be placed entirely above app memory (as on the
new iOS 27.0 embedded VM layout, where debug memory pushes shadow
past kHighMemEnd), there is no need to split shadow into low/high
halves with a middle gap.
- Add kGaplessShadow (Apple-only) to detect this configuration.
- Teach InitializeShadowMemory to reserve one contiguous shadow
region and protect only the shadow-of-shadow when kGaplessShadow
is true, with CHECKs asserting the mapping preconditions.
- Update PrintAddressSpaceLayout to print the single-region layout.
rdar://167657399
[sanitizer_common][Darwin] Add debug memory region support for iOS 27.0
iOS 27.0 bumps the address space from 36 to 39 bits on some devices,
and reserves some address space for sanitizers.
- Add SANITIZER_IOSDEVICE and SANITIZER_EMBEDDED_VM_LAYOUT macros;
bump Darwin iOS/ARM64 SANITIZER_MMAP_RANGE_SIZE from 36 to 39 bits.
- Add ActivateDebugMemory / DebugMemoryActive and tag mmap allocations
above DARWIN_DEBUG_MEMORY_START with VM_MEMORY_DEBUG when using the
debug range; verify returned addresses fall within the expected range.
- Replace GetAppReservedRanges with GetAppRanges, populated from
sysctls on supported devices.
- Extend FindAvailableMemoryRange with a use_debug_vm parameter and
route MapDynamicShadow through it when debug memory is active.
- Centralize mach_vm_region_recurse calls through
internal_mach_vm_region_recurse, which fatals on KERN_DENIED.
rdar://167657399
[BOLT][AArch64] Don't treat entry mapping symbols as function symbols (#230244)
ARM code mapping symbol ($x) at a function entry is treated as a
function alias. When rewriting, it's updated to match the function:
* gets function output size - while mapping syms should be size 0,
* fragment/ICF symbols are derived from it (e.g. $x.cold.0), without
corresponding parent symbol.
Move mapping symbols together with the function but keep 0-sized and
don't add extra symbols for them.
This fixes issues of reading BOLTed binary with such symbols where
perf2bolt/heatmap report:
* "parent function not found for $x.cold.0",
* "owning FILE symbol not found for symbol $x.cold.0".
Test Plan: added split-func-mapping-symbol.s
Assisted-by: Claude Opus 5.5
[CIR][SYCL][NFC] Add sycl-module-id test (#230324)
Port clang/test/CodeGenSYCL/sycl-module-id.cpp to ClangIR.
FWIW I will port the SYCL tests from OG to CIR. That will likely make it
easier to track what is missing.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[CIR][SYCL][NFC] Add functionptr-addrspace test (#230327)
Port clang/test/CodeGenSYCL/functionptr-addrspace.cpp to ClangIR.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[BOLT] Ignore __bolt_reserved syms in perf2bolt/heatmap (#230300)
Aggregation and heatmap don't rewrite the binary, so skipping setting
reserved space is safe. This unblocks aggregation from BOLTed binary
that used and updated these markers.
Test Plan: updated bolt-reserved.test
[WebAssembly] Fall back to SelectionDAG for vector compares in FastISel (#227368)
FastISel's selectICmp and selectFCmp each classify compare operands
with a single scalar test: anything but i64 is treated as i32, and
anything but f64 as f32, vectors included. A vector compare therefore
emits a scalar i32.eq/i32.ne (or f32.eq) that consumes the two v128
registers holding the operands, and the resulting module fails
validation ("type mismatch: expected i32, found v128").
Return false for vector operands in both selectors so the SelectionDAG
lowers them to SIMD compares, and extend the existing
fast-isel-simd128.ll test to cover vector compares on both wasm32 and
wasm64 (UTC-generated checks, with the vectors passed in as arguments).
The vector compare only reaches this path when the rest of the block
does not bail out, which is why it went unnoticed: the common shapes
(a vector icmp feeding extractelement, a bitcast, an intrinsic, or a
branch condition) all miss somewhere in FastISel and the SelectionDAG
revisits the block, replacing the bad compare. A vector compare
[4 lines not shown]
[libc] add sys/inotify (#230314)
The sys/inotify header has linux functions for watching a directory.
This PR adds those functions, and their relevant macros and type. Tests
are fairly simple since the functions are all syscall wrappers.
Assisted-by: Automated tooling, human reviewed
[CIR] Emit calling convention on call sites (#230123)
Adds calling convention on the call sites.
Note: I opted to execlude support for runtime calling convention
(classic getRuntimeCC()) on calls to runtime functions created in CIR
passes (EH, __cxa_atexit, __cxa_guard_*, dynamic_cast, global init).
These keep the default C calling convention and are marked with
TODO(cir) and `MissingFeatures::opFuncCallingConv()` for a follow-up.
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[X86] Emit LSDA call site info for inline asm calls marked `unwind` (#218276)
LLVM currently doesn't consider inline asm calls with the `unwind`
keyword when generating the list of LSDA call sites of a function.
For example, given the following IR
```llvm
declare i32 @rust_eh_personality(i32, i32, i64, ptr, ptr)
declare void @bar()
define void @example() personality ptr @rust_eh_personality {
entry:
call void asm sideeffect alignstack inteldialect unwind "call foo", ""()
invoke void @bar()
to label %cont unwind label %lpad
cont:
ret void
[98 lines not shown]
[BOLT] Keep ambiguous references next to function boundaries valid
A reference into code without a relocation that names its target, e.g. a
RIP-relative LEA whose relocation was against a section symbol, cannot be told
apart from "Next - Delta" and "Prev + Offset" when it lands right before or
after a function start. HHVM built with LLVM 23 has such a reference to
"sqlite3RCStrUnref - 1", which lands in padding and makes BOLT fail with
-strict. Without padding, BOLT silently kept such references relative to the
preceding function. See bolt/test/X86/unanchored-code-reference.s.
BOLT now collects these references from code and, when they are within
--boundary-ref-distance bytes (default 2) of a function start, keeps the
functions around them in place in lite mode. When all functions are processed,
it emits those functions unoptimized and back-to-back with the original bytes
between them, and checks after linking that they kept their size and distance.
Absolute references right before a function start are now relative to that
function. This replaces the workaround for "fptr - 1" in de-virtualized member
function pointer calls, which dropped the relocation and left the input
address in the code.