[lld][WebAssembly] Fix error message when linking wasm64 file with -mwasm32 (#227091)
When `-mwasm32` is explicitly passed and a wasm64 object file is linked,
the error message previously stated:
"wasm32 object file can't be linked in wasm64 mode". Fix this to report
that the wasm64 object file cannot be linked in wasm32 mode.
fix Wasm exceptions + coop threading + shared libraries (#222747)
Prior to this commit, the combination of Wasm exception handling,
cooperative multithreading, and shared libraries was broken.
Specifically, the code generation in `WasmEHPrepare.cpp` involved
direct, cross-library access to `libunwind.so`'s thread-local
`__wasm_lpad_context` variable. However, the ABI used for cooperative
multithreading does not support cross-library access to thread-local
variables.
The solution used here is to add a new `_Unwind_GetWasmLPadContext`
function to `libunwind.so` and use that to get address of the
`__wasm_lpad_context` for the current thread, both in the code generated
by `WasmEHPrepare.cpp` and in the `__gxx_wasm_personality_v0` function
defined in `cxa_personality.cpp`. I've used this strategy
unconditionally for all targets, regardless of whether cooperative
multithreading and/or position-independent are enabled. If desired (e.g.
for performance or code complexity reasons), I could make it conditional
on both of those features being enabled and fall back to using
[2 lines not shown]
[AMDGPU] Fold fpround of fadd and fsub into v_mad/fma_mixlo and mixhi
MadFmaMixFP32Pats turns (fadd x, y) into (fma x, 1.0, y) and (fsub x, y)
into (fma (-y), 1.0, x) so the mix instructions absorb the operation along
with the f16 or bf16 source modifiers. MadFmaMixFP16Pats and
MadFmaMixFP16Pats_t16 only did this for fmul, so a rounded result still
needed a separate convert for a rounding the mix instructions perform
themselves.
Unlike the f32 patterns these do not require an operand to be an fpextend
of an f16, since an fpround on the result always removes the convert. The
rewrite is exact because the mix instructions round the f32 result again
when they write the 16-bit destination, so it stays f32_to_f16(fma(x, 1.0,
y)).
Assisted-by: Claude Code Opus 5
[ProfileData] Only keep module functions when reading ProfileSymbolList
When a sample profile is loaded for a module (SampleProfileLoader), the
profile symbol list is only ever queried for functions of that module:
`PSL->contains(F.getName())` in SampleProfileLoader and
`PSL->contains(CanonFName)` in SampleProfileMatcher. Yet the string-based
reader inserts every symbol of the profiled binary into a DenseSet, in every
compile and every ThinLTO backend that loads the profile.
When the reader has a module, build a small set of that module's function
names (raw and canonical) and only add matching list entries. The list is
still scanned, but nothing outside the module is inserted, so there is no
large hash table to build. Readers without a module (llvm-profdata) still
load the full list. The MD5 symbol list is unaffected.
In a fleet-wide CPU profile of a production clang,
`ProfileSymbolList::read` accounted for 1.7% of all clang cycles and 10% of
ThinLTO backend cycles.
[11 lines not shown]
[gn] bump deployment target to macOS 13 (#227124)
macOS 13 is four years old by now.
This has the effect that lld starts defaulting to chained fixups with
this. Chained fixups reduces `clang --version` from 4 ms to 3.2 ms. (Not
that it matters.)
(Without #227120, chained fixups reduce `clang --version` from 12.5 ms
to 10.9 ms.)
No behavior change.
[ObjC][SEH] Fix clang crash when using finally statements (#176779)
When targeting a platform that does not have funclet-based EH, we push
the finally cleanup (normal edge) and catchall (unwind edge) onto the
EHStack _before_ pushing all catch handlers. The try statement is then
emitted and catch handlers popped from EHStack. Last, the finally
cleanup is popped from EHStack.
For funclet-based EH, we outline and push the finally funclet (of type
`NormalAndEHCleanup`) onto the EHStack _after_ pushing the catch
handlers and never pop it. This results in a crash during codegen when
we try to emit the catch handlers. Not popping the finally cleanup from
the EHStack results in incorrect calls to cleanup handlers in nested
try/catch/finally statements.
I fixed the two issues by:
1. Pushing the finally cleanup first, and
2. Popping it at the end of `CGObjCRuntime::EmitTryCatchStmt`.
Fixes #51899
[gn] bump deployment target to macOS 13 (#227124)
macOS 13 is four years old by now.
This has the effect that lld starts defaulting to chained fixups with
this. Chained fixups reduces `clang --version` from 4 ms to 3.2 ms. (Not
that it matters.)
(Without #227120, chained fixups reduce `clang --version` from 12.5 ms
to 10.9 ms.)
No behavior change.
Add registerToCppTranslation to CppEmmitter.h (#226337)
Currently if one wants to register `mlir-to-cpp` out of tree, they must
include `mlir/InitAllTranslations.h` and depend transitively on all
translation targets (in bazel, `@llvm-project//mlir:AllTranslations`).
This change adds the registration declaration to `CppEmitter.h` so that
one can depend just on the `MLIRTargetCpp` target (or
`@llvm-project//mlir:TargetCpp` in bazel).
This matches the organization of the SMTLib codegen registration in
`mlir/include/mlir/Target/SMTLIB/ExportSMTLIB.h` (though some other
targets like `IRDLToCpp` do it differently).
[gn] Build executables without exported symbols on macOS (#227120)
Speeds up `clang --version` from 12 ms to 4 ms on my system. 8 ms faster
startup isn't a lot, but there's also no reason not to do it.
No intended behavior change.
[profcheck] Exclude find-first-byte-nested.ll (#227121)
We only fixed x86 for LoopIdiom. PR #225576 added a test that looks like
it's just exposing existing propagation issues.
website: Revive donations page
Restore donations links on the project website to the projects donations
page. These links were replaced in the foundation sponsored website
refresh with links to the foundations donations page.
The very first link on this page is to the Foundations donations, but
importantly, it explains where this money actually goes, which has many
common misconceptions, such as the cluster. There are also other links
on the page, such as hardware donations, which we actually need right
now for the aging cluster.
Event: EuroBSDcon 2026
Reviewed by: adrian, bapt, gahr, kevans, obiwac, philip,
Reviewed by: vishwin, wosch,
Reviewed by: Antranig Vartanian <antranigv at freebsd.am>,
Reviewed by: Lukas Engelhardt <lukas.engelhardt at gmx.de>
Differential Revision: https://reviews.freebsd.org/D59505
net/frr9: Install daemons in lib/frr, fix SNMP build and rc.d restart
Daemons move from ${PREFIX}/sbin to ${PREFIX}/lib/frr, matching the Linux
packages and net/frr10. Only vtysh and frr-reload stay in bin/ and sbin/.
This avoids filename collision with net/openbgpd*, net/pimd, net/quagga, etc.
Fix the SNMP build.
Backport frr10's rc.d fixes: a bare "service frr restart" restarted watchfrr
only, leaving the daemons it supervises untouched.
[AMDGPU] Prefer a safe V_PERM_PK16 follower in the scheduler (gfx1250/gfx1251)
Stacked on the post-RA V_PERM_PK16 hazard fixup. V_PERM_PK16 must be
immediately followed by a "safe" instruction (see
SIInstrInfo::isVPermPk16SafeInstr) or the post-RA fixup has to insert a
forced-EXEC V_NOP. Teach GCNHazardRecognizer to bias a safe follower into
the slot right after a V_PERM_PK16 so that V_NOP can be avoided.
Assisted-by: Opus 4.8 Medium
[HLSL] Build elementwise cast results from poison (#225591)
I noticed this unnecessary alloca while doing this pr:
https://github.com/llvm/llvm-project/pull/225519
The change is to initialize vector and matrix elementwise cast results
with poison instead of loading uninitialized temporary storage.
We do this because every result element is overwritten before use,
making the temporary allocation and load unnecessary.