[RISCV] Generalize combineBinOpOfZExt to allow sext (#230472)
This combine narrows binary ops so they're performed at a smaller LMUL.
The combine currently only handles cases where both operands are zext,
but we can also allow sext.
If any operand is sexted then we need to sext the result as the sign bit
may not be zero.
We can't handle udiv + urem if either of the operands are sext, since
sign extending a narrower udiv/urem gives incorrect results. E.g. `udiv
(sext (i8 -1) to i32), (zext (i8 2) to i32)`:
- At i32: `udiv 0xffffffff, 2 = 0x7fffffff`
- Narrowed to i16: `sext (udiv 0xffff, 2) to i32 = sext 0x7fff to i32 =
0x00007ffff`
[RISC-V][MC] Add test for C/Zce implication in .option arch (#230298)
Add MC test coverage for 80f1acc631ca
(https://github.com/llvm/llvm-project/pull/229905). Previously, with
`rv32i` running `.option arch, +zca; .option arch, -zca` resulted in
`error: can't disable zca extension; c extension requires zca extension`,
`.option arch, +zca, +zcb, +zcmp, +zcmt; .option arch, -zcmt` resulted
in `error: can't disable zcmt extension; zce extension requires zcmt extension`,
and `.option arch, +zca, +zclsd; .option arch, +f` resulted in
`error: 'zclsd' and 'zcf' extensions are incompatible`.
This commit was created with the help of AI tools
Pull-Request: https://github.com/llvm/llvm-project/pull/230298
[flang] Classify intrinsic functions as SIMPLE (#223160)
Classify standard intrinsic functions as `SIMPLE` per Fortran 2023 16.1(2).
This change also:
- Classifies extensions as `SIMPLE` where applicable
- Documents the `SIMPLE` and non-`SIMPLE` classifications of extensions
in Extensions.md
- Updates procedure-interface tests affected by the new `SIMPLE`
classification
- Updates F202X.md to reflect the current `SIMPLE` implementation status
Part of the `SIMPLE` procedure support tracked in #221457.
[LLDB] Add user documentation for DIL. (#228275)
Add documentation for users (and programmers) about what DIL is, what it
does, and how it can be controlled.
[ASan][Darwin] Support gapless shadow layout for iOS 27.0 (#217530)
This is the second PR in a series that upstreams support for iOS 27.0
(see #217527 for the first).
When the shadow can be placed entirely above app memory (as on iOS 27),
there is no need to split shadow into low/high halves with a middle gap.
This PR adopts a "gapless" layout in ASAN, in which all of application
memory is in `[kLowMemBeg, kLowMemEnd]` and shadow is `[kLowShadowBeg,
kLowShadowEnd]`. The "high" region is unnecessary and thus unused
(because of the way it is defined, `kHighMemBeg > kHighMemEnd` in this
new layout, which conveniently means that `AddrIsInHighMem(p)` is always
false) -- this is a bit jank and unintuitive, but it keeps the diff
between iOS and other platforms relatively small.
- Add kGaplessShadow (Darwin-only) to detect this configuration.
- Teach `InitializeShadowMemory` to reserve one contiguous shadow region
and protect only the shadow-of-shadow when kGaplessShadow is true
- Update PrintAddressSpaceLayout to print the single-region layout.
rdar://167657399
[sanitizer_common][Darwin] Add debug memory region support for iOS 27.0 (#217527)
Replaces #216885 (but not from a fork so I can use stacked PRs)
iOS 27.0 bumps the address space from 36 to 39 bits on some devices, and
reserves address space for sanitizer shadow memory which is unavailable
to normal applications. This is the first of a series of PRs which
upstreams support for this configuration.
- Add `SANITIZER_IOSDEVICE` and `SANITIZER_EMBEDDED_VM_LAYOUT` macros
for physical devices / devices that support the 39-bit layout; bump
Darwin iOS/ARM64 `SANITIZER_MMAP_RANGE_SIZE` from 36 to 39 bits on
physical devices.
- Add `ActivateDebugMemory` / `DebugMemoryActive` and tag mmap
allocations above `DARWIN_DEBUG_MEMORY_START` with `VM_MEMORY_DEBUG`
when using the debug range; verify returned addresses fall within the
expected range.
- Replace `GetAppReservedRanges` with `GetAppRanges`, populated from
sysctls (on supported devices) which report what ranges are available to
[11 lines not shown]
[AMDGPU] Make named barrier type 1 byte to fix barrier IDs of array elements (#230516)
Since #209746, a pointer in the barrier address space (15) is the
barrier ID, and the backend reads the ID as `ptr & 0x3F`. But
`target("amdgcn.named.barrier", 0)` is still 16 bytes, so GEP to element
`i` of a barrier array adds `16 * i` to the barrier ID. For example, if
`@bars` gets barrier ID 1, `&bars[2]` selects barrier 33 instead of 3.
This PR changes the type to 1 byte, so one array element is one barrier
ID.
Fixes LCOMPILER-2898.
HipStdPar: Don't reprocess math library functions as intrinsics
All uses were already replaced so this only inserted unused nonsense
declarations by replacing the first 4 characters of the function name
with "__hipstdpar". For example acosh would be replaced with
__hipstdparh.
Co-authored-by: Claude <noreply at anthropic.com>
[AMDGPU] Fix missed WMMA C-operand co-exec hazard
The gfx1250 WMMA co-execution hazard check treats only A, B and the
SWMMAC index as registers the in-flight MMA still reads. C (src2 of a
non-SWMMAC WMMA) is missing, so a VALU scheduled into the MMA's shadow
can clobber C and the MMA consumes the new value.
This is latent while C is tied to vdst, since the existing D check then
covers it. It miscompiles where the tie does not hold: for
v_wmma_bf16f32_16x16x32_bf16, whose D is narrower than C, and for the
_threeaddr form of any WMMA.
[OpenMP] Do not reinstall ompd modules in multilibs (#230261)
Summary:
Mulitlibs build the same library with different options, usually placed
into another directory. These libraries do not follow that scheme but
have no real utilitiy in a multilib so just exclude them.
[SLP]Fix crash on split reduction root cast cost
Split roots keep their operands in the combined sub-nodes. The operand
lookup on the root asserts, so leave the context hint unset for them.
Fixes #230407
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/230597
[clang][Sema] Store the evaluated elements of file-scope compound literals (#221390)
Fixes #212106
A file-scope compound literal has to be constant-initialized, and Sema
checks each element in a constant context, where `__builtin_constant_p`
of something it cannot fold simply yields 0. It wrapped each element in
a `ConstantExpr` but never stored the value, so CodeGen evaluated the
element a second time under whatever context it happened to be in. When
the literal's address is taken by a global that needs dynamic
initialization, that context is non-constant, `__builtin_constant_p`
refuses to fold, and `tryEmitGlobalCompoundLiteral` asserts.
Sema now evaluates each element once, as the initializer of a static
object through `EvaluateAsConstantExpr` with a new
`ConstantExprKind::Initializer`, and stores the result in the wrapper it
already creates. CodeGen emits the stored value and no longer
re-evaluates the element. The structural `isConstantInitializer` check
is no longer used for compound literals, and the constant evaluator and
CodeGen now see the same contents for the literal.
RuntimeLibcalls: Require system library members to be libraries
Every SystemRuntimeLibrary now lists only LibcallLibrary and LibraryRef
members, so the inline path that expanded unhomed RuntimeLibcallImpl members
directly into the system's block, including its SystemAvailableImpls
bitset, is dead. Remove it, and error on any member that is not a
library. The generated RuntimeLibcalls.inc is unchanged.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
RuntimeLibcalls: Remove the Default* libcall lists
Every system library now lists only LibcallLibrary members, so nothing
references DefaultRuntimeLibcallImpls. arm64ec's '#'-prefixed set was
still derived from the whole default list minus the width slices and
the Windows exclusions. Build it from the compiler-rt, libm and libc
slices instead, still omitting the calls the Windows runtime lacks.
Take AArch64's fp128 long double calls from the libm slice, and delete
the remaining default and Windows base lists.
The generated tables are unchanged.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
RuntimeLibcalls: Organize functions into libraries
Associate runtime functions with the library which provides them. Add
shared LibcallLibrary defs, like compiler-rt, libm and libc, along with
OS and target specific variants, and replace each target's hand-listed
SystemRuntimeLibrary body with references to them. The resulting libcall
sets are unchanged.
Targets which differ from the shared libraries opt out of individual
functions, e.g. AVR has no sin, cos or sincos, x86 has no fp128 sincosl,
and PPC replaces some f128 compiler-rt helpers. Libraries which should
not merge with the generic variants are Isolated, like arm64ec's
compiler-rt and SPIRV's libc.
Co-authored-by: Claude Opus <noreply at anthropic.com>