[flang][aa] Use Fortran dummy intent in FIR call modref (#227432)
fir::AliasAnalysis::getCallModRef now uses the callee dummy's declared
intent when a pass-by-ref argument aliases the queried variable.
intent(in) is a read, intent(out) is a write, and intent(inout) is both.
A dummy with no visible intent, a callee without a body, or a mismatched
argument count stays a read and a write. The callee is resolved through
the cached symbol table. A local that is not passed is still NoModRef.
[Clang][Sema] Add fortify warnings for fread, fwrite, and fgets (#204337)
Add missing `-Wforitfy-source` diagnostics for `fread`, `fwrite`, and
`fgets`.
This continues the work from #142230.
[flang] Do not honor -fstack-arrays inside offload regions (#227537)
With `-fstack-arrays`, the allocation-placement and stack-arrays passes
move array temporaries from the heap to the stack. Inside an OpenACC
compute construct or a `cuf.kernel` loop, the temporary becomes a
dynamic alloca in the GPU kernel and overflows the device stack:
```fortran
subroutine run()
integer, pointer :: input(:)
integer, allocatable :: output(:)
...
contains
subroutine assign()
!$acc kernels present(input, output)
output(:) = input(:) * 3 ! RHS temporary, dynamic alloca in the kernel
!$acc end kernels
end subroutine
end subroutine
[15 lines not shown]
[NVPTX][TTI] Fix v4i8 scalarization cost (#226071)
Fixes #224808
Handle v4i8 before the generic packed 32-bit vector case so its
dedicated scalarization cost is reachable. Accumulate the cost in the
outer variable and add a cost-model regression test.
---------
Co-authored-by: Justin Fargnoli <jfargnoli at nvidia.com>
[flang][codegen] Report a shape or slice cg-rewrite cannot read (#227625)
cg-rewrite folds a fir.shape, fir.shape_shift, fir.shift or fir.slice
into the code-gen form by reading it through its defining op. A value
that has none cannot be folded. For a slice this went unreported: the
rewrite dropped it and produced a descriptor for the whole array rather
than the section it names. For a shape it reached a cast on a null
defining op.
Report it instead, and say which operand. Rebuilding the value covers a
block argument, but not every case: a slice chosen by an arith.select
has no single value to take apart.
[AMDGPU] Narrow provably-I32 address offsets
Add a `tryNarrowToI32()` function that uses KnowsBits to detect 64-bit
offests that could be in a 32-bit register, allowing us to match
base + offset cases that don't involve literal base + (ext offset)
nodes.
Apply this to global_* instructions and the scalar memory instructions
that take a scalar offset.
This does leave a few dead copies when we have to bail out of the SMEM
handling for creatining a negative offset, but those don't have
end-to-end effect and it looks like you could already get dead MOVs
from that codepath.
AI disclosure: Claude wrote the code, I looked at it.
[AMDGPU] Pre-commit tests for narrowing offsets in address matching
When trying to match either SADDR+VADDR global_* or the s_load_* that
takes a 32-bit offset, we don't check for cases where a 64-bit value
is trivially truncatable to 32 bits. Add tests for these cases.
AI disclosure: Claude wrote these tests
[ObjC] Don't place references to class stubs in __objc_superrefs (#227507)
Loading a reference to a class stub calls objc_loadClassref, which
updates the reference in place. The __objc_superrefs section is a
constant section, so a reference to a class stub cannot be placed there.
Emit such references without a section instead.
rdar://184656330
[flang] Lower enumeration types as named records with a type descriptor
Represent an F2023 enumeration type as !fir.type<...{__ordinal:i32}>
with a real .dt descriptor, instead of a bare i32. Enumeration values
can now be boxed, passed as polymorphic (SELECT TYPE, ALLOCATE, I/O),
and are distinct from INTEGER in descriptors. Ordinary scalar uses
access __ordinal directly through hlfir.designate, with no extra
boxing.
This new approach should address all the review findings thus far.
[BOLT] Support DW_EH_PE_sdata8 encoding in .eh_frame_hdr (#227847)
BOLT always wrote .eh_frame_hdr with 4-byte offsets, truncating those
beyond
2GB. Use 8-byte encoding when needed, as lld does since #179089.
Assisted-By: Opus 5.5
[lld][AArch64] Support R_AARCH64_TLSLE_LDST*_TPREL_LO12 (#227629)
Follow-up to #227173. Support the non-NC local-exec TLS load/store
relocations, emitted for `ldr/str xN, [xN, :tprel_lo12:sym]` (previously
"unknown relocation"). They resolve like the `_NC` variants — imm12
holds bits 11:scale of the TP offset — plus an unsigned 12-bit range
check on the full offset. Boundaries and encodings verified identical to
GNU ld 2.46.
Testing: new `aarch64-tls-le-ldst.s` (encodings, out-of-range boundary,
alignment); extended `aarch64-tls-le.s`; full `lld/test/ELF` passes.