[Verifier] Strengthen type constraint (#229880)
When I asked the LLM to hack llubi, it created some inputs with an
invalid intrinsic like llvm.vscale.v2i64, bypassing the verifier. This
patch uses predicates introduced by
https://github.com/llvm/llvm-project/pull/204138 to specify
scalar-only/vector-only overload types.
Some redundant checks in Verifier.cpp are simplified. There are two
kinds of additional constraints in Verifier.cpp, as they cannot be
represented by existing predicates in tablegen:
1. matrix intrinsics now only accept fixed vectors.
2. vector.partial.reduce requires the accumulator vector and input
vector to share the same element type.
I have checked that the rejected form is either rejected by llc/opt
(e.g., instruction selection failures or direct use of
`cast<VectorType>`) or explicitly specified in LangRef.
Aided by DeepSeek-V4.1-Flash.
AMDGPU: Use functions instead of kernels in f16 operation tests (#229899)
These tests used kernels with loads from pointer arguments to supply
operand values, relying on -amdgpu-scalarize-global-loads=false to keep
the loaded values in VGPRs. Use function arguments instead.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[LLVMABI] Classify types at their ABI size (#227889)
The x86-64 classifier compared value widths against Clang type sizes, so
types with padding were misclassified. For example, `struct { long
double __attribute__((vector_size(16))) v; }` coerced to `<2 x double>`
instead of `<1 x x86_fp80>`.
getABISizeInBits() now gives every type the size Clang's getTypeSize()
gives it, and the classifier uses it wherever Clang uses getTypeSize().
This also applies to _BitInt and bool arrays and bit-fields, vectors
with a non-power-of-two element count, bool vectors, and vectors of
sub-byte _BitInt. isIllegalVectorType now sends only __int128 vectors to
memory, not _BitInt(128) ones, as Clang does. isSingleElementStruct is
shared, so AMDGPU picks up the fix too.
Assisted-by: Claude Code / Claude Opus 5.5
[CIR] Fix 'vague' linkage globals. (#228586)
Namespace scope globals with 'vague' (internal/linkonce ODR) need to be
'guarded' to prevent them from being initialized in multiple TUs, which
typically shows itself with multiple 'free' operations happening (since
it is registered to 'free' multiple times: the initial allocation just
leaks).
This patch gets that correct. Additionally, these have the idea of an
'associated variable' when added to teh Global Ctor/Dtor, so this
modifies that to make sure that is correct.
As this infrastructure reuses the
'static_local_guard'/'static_local_info', I've generalized that naming
with the help of Claude who hopefully isn't pulling a fast one on us :)
net/dnscap: Update dnscap from version 1.4.1 (from 2015) to 2.5.1
Prompted by jperkin's MacOS 27 bulk build results
10+ years of changes are too many to summarise here, but TL;DR is that
+ there are a lot more dependencies (on a lot of archivers/compression
libraries), openssl, and ldns,
+ the package name in pkgsrc is now dnscap2,
+ and dnscap now supports plugins (${PREFIX}/bin/dnscap-rssm-rssac002
is one such).
[Clang] Diagnose incompatible `weak` and `ifunc` attributes (#229393)
Clang previously accepted both `weak` and `ifunc` attributes on the same
function, but silently ignored `weak` when emitting the IFUNC. Since
`ifunc` functions must not have weak linkage, the new mutual exclusion
rule diagnoses the conflict during attribute processing and
redeclaration merging.
Using `#pragma weak` after an `ifunc` declaration adds an implicit
`weak` attribute directly, bypassing the normal mutual exclusion checks.
The pragma handler now checks attribute compatibility as well. If the
pragma follows a redeclaration, the check uses the original IFUNC
definition, because the redeclaration does not inherit the `ifunc`
attribute.
Link:
https://github.com/llvm/llvm-project/issues/220923#issuecomment-5993507925
Fixes #220923
[SimplifyLibCalls] Fix profiles in optimizeFFS (#229842)
We were creating a select with a condition that compares the input value
to the null value for that type. Mark that condition as being unlikely
as that should hold for most callsites in most applications.
[AMDGPU] Allow commuting immediates out of src0 when legal
isLegalToSwap refused to move any non-inline constant out of src0, so
commuting an instruction with a literal in src1 could not be undone.
AMDGPULowerVGPREncoding relies on undoing it and, on gfx1250, either hit
"Failed to restore commuted instruction" or kept the commuted
instruction with the wrong VGPR MSB mode, which made it read the wrong
VGPRs.
Allow an immediate to leave src0 when the other operand can hold it.
VOPD formation now accepts a V_DOT2 with a literal in src1 if
isLegalToSwap allows the swap, and commutes it when the pair is built,
so those pairs are still formed.
Co-Authored-By: Claude <noreply at anthropic.com>
[CAS][unittests] Invalid ID should not assert but return error (#229892)
`llcas_digest_parse` implementation in CASPluginTest plugin doesn't
follow the contract of the API that invalid ID should return error
instead of assert.
[KnownFPClass] Correct denormal handling for `KnownFPClass::roundToIntegral` (#219700)
Fixes https://github.com/llvm/llvm-project/issues/217412
Part of https://github.com/llvm/llvm-project/pull/219091
Fixes `KnownFPClass::roundToIntegral` denormal handling, particularly
around DAPZ (Denormals are positive zero).
Some sign preserving deductions were also improved slightly.
I also completed the TODO for `LLT FPInfo`.
AI Disclosure:
I used ChatGPT Codex (sol 5.6) to help write the tests.
[NFC][SLP] Refactor to capture repeated pattern in 'TreeEntry::isReverseOrder()' (#229890)
Pattern is repeated a lot, seems good to clean up into its own function.
AI Usage: Assisted by Codex
[ThinLTO] Only replace uses in ExportM when promoting CFI functions (#229634)
In `promoteInternals`, we only need to replace uses of `ExportGV` with
`ExternalAlias` in `ExportM` when `MustPromote` is true. Otherwise, we
also do it for globals that reference themselves in their own
initializer (like relative vtables), which then causes lowering
problems. For example, relative vtables, which compute
`dso_local_equivalent @fn - @vtable`, replacing `@vtable` with the alias
breaks `AsmPrinter` lowering of `dso_local_equivalent`, which expects
the subtraction base to be the global being emitted, not an alias.
addresses http://crbug.com/570039617
fixes regression introduced in PR #225173
Co-authored-by: Nico Weber <thakis at chromium.org>
[AMDGPU] Mark 64-bit VOPC as DP-MACC
A let around a multiclass does not reach records the multiclass
inherits from another multiclass, so IsDPMACCInstruction was never
set on 64-bit V_CMP/V_CMPX. In expert scheduling mode, overwriting a
compare's source after a later CS-MACC VALU waited for va_vdst(1)
instead of va_vdst(0).
Set the flag inside the multiclass bodies and resolve the
V_CMP(X)_CLASS_F64 FIXMEs.
Change-Id: I3b08ca78cdc3455ff91d085ee04a39f23034cda9
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[AMDGPU] Add more cases to the "this is a wave ID" recognizer (#177713)
Previously, we'd only catch things that were actually wave IDs
(workitem_id_x >> log2(wavefrontsize) or ands with a mask that cleared
off the within-lane bits ... and all that only in cases where we had
static assurances about the workgroup size) when the operation was being
performed directly on the ID. This commit adds a test for the mask case
and also extends the matcher to support
1. The intrinsic being cast to another integer type before being shifted
2. A mask being applied to the intristic before the shift occurs, so
that the computation naturally written as (x % 128) / 64 which optimizes
to (x & 0x7f) >> 6 is properly recognized as wave-uniform.
I've noticed that there's a higher-level problem here. These sorts of
quasi-wave-IDs are often needed to compute, for example, the address for
an LDS DMA on gfx950 or will be used on gfx1250 when computing parts of
TDM descriptors. These values need to be in SGPRs, but the VGPR-ness of
the workitem ID causes all the address computation to be forced into
[2 lines not shown]
graphics/converseen: Update to 0.15.2.9
ChangeLog: https://converseen.fasterland.net/changelog/
* Fixed potential DLL conflicts on Windows by restricting the DLL search path
before initializing ImageMagick
* Various Bugfixes
AMDGPU: Use functions instead of kernels in f16 operation tests
These tests used kernels with loads from pointer arguments to supply
operand values, relying on -amdgpu-scalarize-global-loads=false to keep
the loaded values in VGPRs. Use function arguments instead.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
cad/openvsp: Update to 3.54.0
ChangeLog:
https://openvsp.org/blogs/announcements/2026/10/06/openvsp-3-54-0-released
Features:
* Clone Geom
* Flip for every Geom
* Split and stitched STEP and IGES export (File -> Export)
* POGS output from CFDMesh
* CFDMesh faster single threaded
* CFDMesh multi-threaded
* CFDMesh more robust
* CFDMesh higher quality
* STEP/IGES Export greatly improved
* Hinges as all-moving control surfaces for CBAERO via CFDMesh
* Design variables have user determined bounds
[17 lines not shown]