[mlir][OpenACC] Record gang(static:) chunk size on lowered loops (#227037)
Lower OpenACC `gang(static:)` through compute lowering as
`acc.chunk_size`. A constant size is stored directly; `static:*` and a
non-constant size are recorded as -1. `gang(static:)` also maps to gang
dimension 1.
[gn] Fix Interpreter/emulated-tls.cpp on mac (#227132)
In, #225475 (5dfe8605911e), CLANG_HAVE_EMUTLS_GET_ADDRESS got hardcoded
to 0 in the GN build, while the CMake build has a config-time check for
__emutls_get_address.
On mac and linux, compiler-rt provides that symbol, so set
CLANG_HAVE_EMUTLS_GET_ADDRESS to 1 there. That way, __emutls_get_address
is force-linked as absolute symbol and things are happy.
(With it set to 0. clang-repl dlsym()s for the symbol, which works on
Linux where it's found in libgcc_s.so.1, but it doesn't work on mac.)
[flang][Semantics] Warn on BIND(C) interfaces with assumed-shape/rank dummies (#225965)
## Motivation
Fortran 2018 §18.3.6 requires assumed-shape, deferred-shape, and
assumed-rank `BIND(C)` dummy arguments to be passed using a CFI
descriptor (`CFI_cdesc_t`). Some existing C/C++ interfaces predate this
requirement and expect such an argument to be passed by bare address
instead. Declaring a `BIND(C)` Fortran interface against such code
currently produces a silent calling-convention mismatch — no diagnostic
warns the user that the ABI they've declared doesn't match what
pre-2018 C/C++ code expects.
## What this PR does
Adds a portability warning in `CheckSubprogram`
(`flang/lib/Semantics/check-declarations.cpp`) that fires when a
`BIND(C)` interface declares an assumed-shape or assumed-rank dummy
argument, under a new `UsageWarning` category, `BindCArrayDescriptor`.
[40 lines not shown]
MIR: Serialize MachineBasicBlock::MaxBytesForAlignment
Fix missing serialization of another field. The alignment was
already handled. The name is a bit verbose. Some places call it
"MaxSkip" which matches the name of the 2nd operand to the .p2align
directive this corresponds to.
Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
[libc++] Move test tools CI job back to k8s runner sets (#226564)
While the issue with k8s runners in #226230 is still not resolved, it is
better to target the k8s runners and get transient failures than to
target the old runners and have the jobs hang forever (there seems to be
no runners registered in llvm-premerge-libcxx-runners).
Co-authored-by: Aiden Grossman <aidengrossman at google.com>
[lld][WebAssembly] Fix error message when linking wasm64 file with -mwasm32 (#227091)
When `-mwasm32` is explicitly passed and a wasm64 object file is linked,
the error message previously stated:
"wasm32 object file can't be linked in wasm64 mode". Fix this to report
that the wasm64 object file cannot be linked in wasm32 mode.
fix Wasm exceptions + coop threading + shared libraries (#222747)
Prior to this commit, the combination of Wasm exception handling,
cooperative multithreading, and shared libraries was broken.
Specifically, the code generation in `WasmEHPrepare.cpp` involved
direct, cross-library access to `libunwind.so`'s thread-local
`__wasm_lpad_context` variable. However, the ABI used for cooperative
multithreading does not support cross-library access to thread-local
variables.
The solution used here is to add a new `_Unwind_GetWasmLPadContext`
function to `libunwind.so` and use that to get address of the
`__wasm_lpad_context` for the current thread, both in the code generated
by `WasmEHPrepare.cpp` and in the `__gxx_wasm_personality_v0` function
defined in `cxa_personality.cpp`. I've used this strategy
unconditionally for all targets, regardless of whether cooperative
multithreading and/or position-independent are enabled. If desired (e.g.
for performance or code complexity reasons), I could make it conditional
on both of those features being enabled and fall back to using
[2 lines not shown]
[AMDGPU] Fold fpround of fadd and fsub into v_mad/fma_mixlo and mixhi
MadFmaMixFP32Pats turns (fadd x, y) into (fma x, 1.0, y) and (fsub x, y)
into (fma (-y), 1.0, x) so the mix instructions absorb the operation along
with the f16 or bf16 source modifiers. MadFmaMixFP16Pats and
MadFmaMixFP16Pats_t16 only did this for fmul, so a rounded result still
needed a separate convert for a rounding the mix instructions perform
themselves.
Unlike the f32 patterns these do not require an operand to be an fpextend
of an f16, since an fpround on the result always removes the convert. The
rewrite is exact because the mix instructions round the f32 result again
when they write the 16-bit destination, so it stays f32_to_f16(fma(x, 1.0,
y)).
Assisted-by: Claude Code Opus 5
[ProfileData] Only keep module functions when reading ProfileSymbolList
When a sample profile is loaded for a module (SampleProfileLoader), the
profile symbol list is only ever queried for functions of that module:
`PSL->contains(F.getName())` in SampleProfileLoader and
`PSL->contains(CanonFName)` in SampleProfileMatcher. Yet the string-based
reader inserts every symbol of the profiled binary into a DenseSet, in every
compile and every ThinLTO backend that loads the profile.
When the reader has a module, build a small set of that module's function
names (raw and canonical) and only add matching list entries. The list is
still scanned, but nothing outside the module is inserted, so there is no
large hash table to build. Readers without a module (llvm-profdata) still
load the full list. The MD5 symbol list is unaffected.
In a fleet-wide CPU profile of a production clang,
`ProfileSymbolList::read` accounted for 1.7% of all clang cycles and 10% of
ThinLTO backend cycles.
[11 lines not shown]
[gn] bump deployment target to macOS 13 (#227124)
macOS 13 is four years old by now.
This has the effect that lld starts defaulting to chained fixups with
this. Chained fixups reduces `clang --version` from 4 ms to 3.2 ms. (Not
that it matters.)
(Without #227120, chained fixups reduce `clang --version` from 12.5 ms
to 10.9 ms.)
No behavior change.
[ObjC][SEH] Fix clang crash when using finally statements (#176779)
When targeting a platform that does not have funclet-based EH, we push
the finally cleanup (normal edge) and catchall (unwind edge) onto the
EHStack _before_ pushing all catch handlers. The try statement is then
emitted and catch handlers popped from EHStack. Last, the finally
cleanup is popped from EHStack.
For funclet-based EH, we outline and push the finally funclet (of type
`NormalAndEHCleanup`) onto the EHStack _after_ pushing the catch
handlers and never pop it. This results in a crash during codegen when
we try to emit the catch handlers. Not popping the finally cleanup from
the EHStack results in incorrect calls to cleanup handlers in nested
try/catch/finally statements.
I fixed the two issues by:
1. Pushing the finally cleanup first, and
2. Popping it at the end of `CGObjCRuntime::EmitTryCatchStmt`.
Fixes #51899
[gn] bump deployment target to macOS 13 (#227124)
macOS 13 is four years old by now.
This has the effect that lld starts defaulting to chained fixups with
this. Chained fixups reduces `clang --version` from 4 ms to 3.2 ms. (Not
that it matters.)
(Without #227120, chained fixups reduce `clang --version` from 12.5 ms
to 10.9 ms.)
No behavior change.
Add registerToCppTranslation to CppEmmitter.h (#226337)
Currently if one wants to register `mlir-to-cpp` out of tree, they must
include `mlir/InitAllTranslations.h` and depend transitively on all
translation targets (in bazel, `@llvm-project//mlir:AllTranslations`).
This change adds the registration declaration to `CppEmitter.h` so that
one can depend just on the `MLIRTargetCpp` target (or
`@llvm-project//mlir:TargetCpp` in bazel).
This matches the organization of the SMTLib codegen registration in
`mlir/include/mlir/Target/SMTLIB/ExportSMTLIB.h` (though some other
targets like `IRDLToCpp` do it differently).
[gn] Build executables without exported symbols on macOS (#227120)
Speeds up `clang --version` from 12 ms to 4 ms on my system. 8 ms faster
startup isn't a lot, but there's also no reason not to do it.
No intended behavior change.
[profcheck] Exclude find-first-byte-nested.ll (#227121)
We only fixed x86 for LoopIdiom. PR #225576 added a test that looks like
it's just exposing existing propagation issues.