[offload][omp] Mark shlib_global_var test unsupported for NVIDIA (#227098)
The shared library test introduced in #226980 seems to fail with NVIDIA
backend. Marking it unsupported.
[Verifier] Validate !tbaa.struct metadata (#225910)
Check that !tbaa.struct operands come in (offset, size, tag) triples
with constant offset and size.
[ADT] Remove CRTP from DenseMapBase (NFC) (#227063)
This patch removes CRTP from DenseMapBase by replacing DerivedT with
StorageT (DenseMapStorage or SmallDenseMapStorage).
DenseMapBase now owns the Storage member by composition and provides
the entire user-facing map interface -- from constructors, the
destructor, and operator= to find, try_emplace, and erase. DenseMap
and SmallDenseMap simply specialize DenseMapBase with their respective
storage types.
This completes the effort to replace CRTP in DenseMapBase with
composition (see #168255, #226664, and #226882).
Assisted-by: Antigravity
devel/kf5-knotifications: Disable DBusMenu support
libdbusmenu-qt is not needed for Qt6 apps, the project is dead, the
port will be removed.
While here drop unused dependencies.
PR: 298946
[InstCombine] Match swapped form of truncating saturation clamp (#226614)
Extend the fold added in #189703 to also handle the inverted select:
trunc (select (icmp ugt A, DestTy_umax), sext(icmp sgt A, 0), A) -->
trunc (smin (smax (0, A), DestTy_umax))
InstCombine canonicalizes (A & NegPow2) != 0 into the ult form with
swapped select operands, but if SCCP first rewrites the compare as
icmp uge A, C, InstCombine only turns it into icmp ugt and never swaps
the select, so the original fold is missed.
While here, match the compare constant with m_APInt instead of
m_Constant + getUniqueInteger, and build TruncatedMax directly with
APInt::getLowBitsSet. Comparing the constant against TruncatedMax + 1
or TruncatedMax makes the separate zero check unnecessary.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[CIR] Correct 2 lowering bugs of atomic cmp-xchng builtins (#227054)
This patch fixese two bugs that showed up in a benchmark.
First; convertToAtomicIntPointer was zero-filling the source object
directly, rather than the temporary. The result was that anything that
would not be overwritten thanks to the power-of-2 write, would be
incorrect, and corrupted.
Second; emitAtomicCmpXchg didn't set the 'old' value back into the real
object. This ends up doing an additional argument on this function that
better matches classic-codegen.
[libc] Add getpwent_r entrypoint (#226966)
Add the reentrant password database iteration entrypoint getpwent_r.
getpwent_r reads the next password database record from the stream into
the caller-supplied struct passwd and buffer, returning 0 on success,
ENOENT at end-of-file, or an error number (such as ERANGE) on failure.
When a record exceeds the caller buffer size, FlatFileDatabase::getnext
rewinds the stream to the beginning of that record so that a subsequent
retry with a larger buffer reads the same entry.
* Add getpwent_r entrypoint and pwd::read_next fixed-buffer overload
* Rewind stream on ERANGE in FlatFileDatabase::getnext(EntryType *,
span<char>)
* Define getpwent_r in include/pwd.yaml and Linux entrypoints.txt
* Add unit tests for getpwent_r
Assisted-by: Automated tooling, human reviewed.
[CIR] Fix order of creation so that lit test will not fail (#226706)
This patch fixes the order of creation, otherwise the compiler may
evaluate one before the other and the lit test fail.
FastISel: Assert the emitted instruction defines the result
The fallback path copied the result out of implicit_defs()[0], assuming
the first implicit physical register def is the result. That is an X86
assumption about MUL/IMUL, and it is unreachable for all but
fastEmitInst_r: FastISelEmitter skips any instruction whose first
operand is not an output register, so every opcode reaching these
helpers from generated code has an explicit def.
Co-Authored-By: Claude Opus 5 <noreply at anthropic.com>
[SDPatternMatch] Make m_SetCC work like m_ICmp from IR PatternMatch. (#226623)
The condition code is stored an operand, but we don't need to expose
that to the interface.
This adds 2 signatures of m_Setcc, one that takes 2 operands and matches
any condition code and one that takes the matched condition code by
reference. For m_SpecificCondCode cases, I've added m_SpecificSetCC.
Similar changes have been applied to m_SelectCC and m_SelectCCLike.
Out of tree targets will need to update to the new interface.
Assisted-by: Claude
hwpmc_amd: add PerfMonV2 global-control path
Add support for AMD PerfMonV2 (Family 19h+) core counters, which need
both the per-counter EVSEL enable bit and the global GLOBAL_CTL bit set
to count. Detects PerfMonV2 at init and switches to v2-specific
start/stop/interrupt handlers; older CPUs and L3/DF counters keep using
the classic path unchanged.
Adds a read-only sysctl, kern.hwpmc.amd_perfmon_v2, to report which path
is active.
Signed-off-by: Andre Silva <andasilv at amd.com>
Reviewed by: Ali Mashtizadeh <ali at mashtizadeh.com>
Sponsored by: AMD
Differential Revision: https://reviews.freebsd.org/D58256
[mlir][OpenACC] Materialize routine bind targets before parallel call rewriting. (#227035)
Separate string-bound OpenACC routine target materialization from call
rewriting.
Add a module-level pass that creates declarations for string-bound
routine targets before the existing per-function ACCBindRoutine pass
rewrites calls. Run the new pass once on the outer module and once on
each GPU module. Remove declaration creation from ACCBindRoutine.
Extend coverage with two outer-module callers and two GPU callers that
reference the same bound routine, a GPU-module-local routine case, and
negative tests for missing materialization and invalid pass placement.
Update the pipeline test to verify ordering in both outer-module and
GPU-module pipelines.
[ARM] Replace uses of ARM::NoRegister with Register() or isValid() NFC (#224083)
Same as https://github.com/llvm/llvm-project/pull/220227 but for the ARM
backend. This replaces remaining uses of ARM::NoRegister and unsigned
with the Register class in lib/Target/ARM.
Add Centaur cpuid and family 7 to check whether the TSC counter is invariant.
Recent Zhaoxin CPUs have an invariant TSC and advertise that through
CPUID 80000007, as Intel and AMD do.
Allows choosing the TSC timecounter as more reliable than HPET.
XXX pullup-11, and perhaps netbsd-10.
[LoopIdiomVectorize] Don't add the match-index block to the parent loop when it exits (#225576)
`expandFindFirstByte` unconditionally adds BB4, the block that computes
the index of the match, to the parent loop. BB4 branches only to
ExitSucc. When the `find_first_of` idiom is nested inside another loop,
and a match exits that enclosing loop, BB4 always leaves the parent loop
and so is not part of it. `LoopInfo` is then inconsistent, and
`LoopBase::verifyLoop()` fails with `"Loop block has no in-loop
successors!"`.
A release build compiles that check out. It instead crashes later, in a
pass that consumes the stale analysis. IndVarSimplify is the one seen in
practice.
Only add BB4 to the parent loop when the parent loop contains ExitSucc.
The sibling `expandFindMismatch` is unaffected, because its success path
lands in a block split out of the preheader, which is inside the parent
loop already.
[6 lines not shown]