[flang] Lower enumeration types as named records with a type descriptor
Represent an F2023 enumeration type as !fir.type<...{__ordinal:i32}>
with a real .dt descriptor, instead of a bare i32. Enumeration values
can now be boxed, passed as polymorphic (SELECT TYPE, ALLOCATE, I/O),
and are distinct from INTEGER in descriptors. Ordinary scalar uses
access __ordinal directly through hlfir.designate, with no extra
boxing.
This new approach should address all the review findings thus far.
[BOLT] Support DW_EH_PE_sdata8 encoding in .eh_frame_hdr (#227847)
BOLT always wrote .eh_frame_hdr with 4-byte offsets, truncating those
beyond
2GB. Use 8-byte encoding when needed, as lld does since #179089.
Assisted-By: Opus 5.5
[lld][AArch64] Support R_AARCH64_TLSLE_LDST*_TPREL_LO12 (#227629)
Follow-up to #227173. Support the non-NC local-exec TLS load/store
relocations, emitted for `ldr/str xN, [xN, :tprel_lo12:sym]` (previously
"unknown relocation"). They resolve like the `_NC` variants — imm12
holds bits 11:scale of the TP offset — plus an unsigned 12-bit range
check on the full offset. Boundaries and encodings verified identical to
GNU ld 2.46.
Testing: new `aarch64-tls-le-ldst.s` (encodings, out-of-range boundary,
alignment); extended `aarch64-tls-le.s`; full `lld/test/ELF` passes.
[clang][OpenMP] Don't use fused dist schedule for teams loop emitted as distribute
Fix teams loop reductions lowered as 'distribute' lose their loop.
Claude assisted with this patch.
[lldb][Windows] Fix import of NtQueryInformationThread (#227471)
`NtQueryInformationThread` is from `ntdll.dll`, not `Kernel32.dll`. We
never tested this. I'll add a test in the next PR when testing the
gdb-remote handler.
[mlir][Vector] Reject scalable reduction in warp distribution (#225267)
`WarpOpReduction` distributes a `vector.reduction` across warp lanes by
computing `numElements = vectorType.getShape()[0] / warpSize` and
building the per-lane type as a plain `VectorType::get({numElements},
...)`, with no check for a scalable operand.
For a scalable reduction vector, e.g. `vector<[32]xf32>` with warp size
32, this silently drops the scalable marker: the per-lane type becomes
the fixed `vector<1xf32>` instead of `vector<[1]xf32>`, so only 32 total
elements get reduced across lanes instead of `32*vscale`.
Reject a scalable reduction operand before any IR is created, and add a
negative test.
NVPTX: Drop LiveVariables from the register allocation pipeline (#225183)
The optimized RegAlloc pipeline ran LiveVariables only to satisfy
PHIElimination and TwoAddressInstruction, both of which no longer need it.
Remove the LiveVariables run (and, in the new pass manager, the
UnreachableMachineBlockElim that was there only as a LiveVariables
prerequisite).
Co-authored-by: Claude (Claude-Opus-4.8)
[LAA] Add tests with loops with may-not-return calls (NFC) (#228102)
Add tests for runtime-check bounds of loops containing a call that may
not return.
[clang-repl] Flush CUDA device bootstrap module before the first PTU (#226975)
The host path sets the bootstrap module aside with CacheCodeGenModule()
once the initial action has run, so the first PTU starts from a fresh
module. The CUDA device path skipped this step. The device module that
HandleTranslationUnit had already finalized during bootstrap stayed
current, and the first device PTU finalized it a second time, resulting
in CodeGen adding every module flag twice. The IR verifier then rejects
the module.
To reproduce, on a build with assertions, or with
`-fverify-intermediate-code` (hidden on release builds since the driver
disables the verifier there), the first input to `clang-repl --cuda`
fails:
```
module flag identifiers must be unique (or of 'require' type)
!"nvvm-reflect-ftz"
module flag identifiers must be unique (or of 'require' type)
[12 lines not shown]
CodeGen: Drop the LiveVariables parameter from convertToThreeAddress (#225182)
This was used for analysis updates, but now the analysis is being removed.
Co-authored-by: Claude (Claude-Opus-4.8)
[lldb][Windows] Move extended info to host thread (#227470)
I want to add support for reading thread locals on Windows. To do this,
we need to know the address of the TEB. This was implemented for
`TargetThreadWindows`, but when using lldb-server, we didn't have this
info. As a first step, move this to the `HostThreadWindows`, so both the
target thread and native thread can call it.
[DataFlowSanitizer] Properly add ext attributes on arguments as needed. (#225443)
TargetLibraryInfo is used to compute the extension attributes, in part by a new
getExtAttrForI8Param() method. It currently always returns ZExt (or SExt) but is
used so that a target can easily override this if needed.
[CodeGen] Remove unused SpillPlacement::Linked (NFC) (#227989)
The last use was removed on May 19, 2016 in commit
b926bdac4c18e0f31d827dec482f207856e88e1e.
Assisted-by: Antigravity
[OpenMP] Fix missing implicit barriers in device worksharing loops (#227735)
Device worksharing lowering discarded `NeedsBarrier`, omitting implicit
barriers after loops without `nowait`. Forward the requirement and emit
the barrier outside the outlined loop body, ensuring all participating
threads synchronise.
Co-authored-by: Codex <codex at openai.com>
[LAA] Add stencil group merging to reduce runtime pointer checks (#187252)
Take this loop, where S is only known at runtime:
for (i = 0; i < N; i++)
Out[i] = In[i - S] + In[i - 1] + In[i] + In[i + 1] + In[i + S];
groupChecks puts In[i - 1], In[i] and In[i + 1] in one group, because
their bounds differ by a constant. In[i - S] and In[i + S] differ from
the rest by a multiple of S, so each stays in its own group. That is
3 groups for In and 3 checks against Out. With two or three strides,
as in 3D stencils, the count grows fast. The Einstein Toolkit / Cactus
CCZ4 code has loops with thousands of checks, and the vectorizer gives
up on them.
This patch adds mergeStencilGroups, which runs after groupChecks. For
the loop above it makes one group for In, from In[i - S] to In[i + S].
One check against Out is enough, plus a check that 1 <= S <= Max.
[27 lines not shown]
[Offload][L0][NFC] Remove old ELF format support (#228011)
After we switched to using Offload Binary (for OpenMP) or direct SPIR-V
images (for SYCL) this support is not used anymore.
Assisted by Claude.
[TableGen] Allow AsmWriter to generate uint64_t tables. NFC. (#227641)
Each OpInfo entry in the generated AMDGPUInstPrinter::getMnemonic
carries 8 bytes of data. Previously GenAsmWriter would split that into
two uint32_t tables for no good reason. Generating a single uint64_t
table makes for shorter output and slightly better generated code.
[AMDGPU][SROA] Expand cast chain handling to floating point types
CreateBitPreservingCastChain does not produce inttoptr or ptrtoint for floating-point types. When promoting structs like { float, float } to <2 x float>, this can lead to pointers being bitcast directory to <2 x float> which is invalid. This change expands the use of the intermediate to these cases