[Hexagon] Add v79 QFloat HVX runtime conversion funcs (#229481)
Add four assembly files for half-float to integer conversions using
QFloat (qf16) on Hexagon V79.
Each function provides IEEE round-to-nearest-even or fast
round-half-away-from-zero conversion for signed and unsigned 8-bit and
16-bit integer destinations.
V79 lacks direct half-float to integer conversion. V75 and below have
native IEEE vcvt instructions,
while V81 and newer have dedicated hardware support.
Co-authored-by: Sumanth Gundapaneni <sgundapa at quicinc.com>
[InstCombine] Don't narrow FP ops when the destination range is too small (#230379)
Fixes #229651.
InstCombine narrows `fptrunc (binop (fpext X), (fpext Y))` into a binop
in the destination type, but it only compared the significand widths.
`bfloat` has fewer significand bits than `half` but a much larger
exponent range, so `bfloat` operands were narrowed to `half` and large
values overflowed to infinity:
```llvm
%x = fpext bfloat %a to double
%y = fpext bfloat %b to double
%sum = fadd double %x, %y
%r = fptrunc double %sum to half
```
For `65536 + -65536` this returns `+0.0`, but after narrowing both
operands become `inf` and `-inf`, and the sum is NaN.
[4 lines not shown]
[AMDGPU] NFC: Drop constexpr from getFlavorName and getCoExecMask (#230334)
Both functions end in llvm_unreachable. GCC 8 rejects a constexpr
function that can reach a call to llvm_unreachable_internal when
assertions are enabled, which breaks the clang-ppc64le-linux-test-suite
and clang-ppc64le-linux-multistage builders.
#203603 already fixed this for getFlavorName, but the constexpr came
back when the function moved into AMDGPUCoExecInfo.h in #204077. Neither
function is used in a constant expression.
[LLVMABI][AARCH64] Implement Pure Scalable Type handling (#227504)
This change adds support for classifying Pure Scalable Type arguments
for AArch64 targets in the LLVM ABI library. Pure Scalable Types are
passed in registers, expanded if necessary, unless the argument is
unnamed or there are not sufficient registers available, in which case
they are passed indirectly. Pure Scalable Types are treated as single
named arguments when used as return types.
Assisted-by: Cursor / various models
[Hexagon][NFC] Fix duplicate A2_tfr check in HexagonGlobalScheduler (#227762)
Fix a typo where both sides of a logical AND checked the same opcode
(A2_tfr). The second check should be A2_tfrsi.
[RISCV] Correct the CFA offsets for the stack probe loop. (#230322)
We need to take into account that we may have already done a
FirstSPAdjust. This is the same bug as #164805, which was fixed
for the unrolled probes in #166616 but not for the probe loop.
Fixes #230291.
Co-Authored-By: Claude Opus 5.5 <noreply at anthropic.com>
[Bazel][llvm] Generate InlinerModels and RegAllocEvictModels headers (#230576)
Fixes build failures following commit c90b6414e607 ([mlgo] Allow passing
pre-emitc-ed models (#227941)) where InlinerModels.h and
RegAllocEvictModels.h are unconditionally included in MLInlineAdvisor
and MLRegAllocEvictAdvisor.
LLM-aided
[AArch64] Add NVCAST to inputs of AArch64ISD::REV32/REV64 to match SDTypeProfile. (#230211)
The SDTypeProfile says the input type should match the result type. Add
NVCASTs to satisfy this constraint.
Found by adding SDTCisSameAs to SDNodeInfo::verifyNode.
[mlir][x86] Fix VNNI operand load offsets (#230205)
Fixes the tile load offsets of VNNI operands in the AMX contraction
lowering.
An offset into the VNNI packed dims was not scaled by the VNNI factor,
so loads at a non-zero offset read the wrong data. VNNI operands whose
buffer's innermost dim is not the static VNNI factor are now rejected,
as their tiles cannot be loaded directly.
Assisted-by: Claude
NAS-144385 / 28.0.0-BETA.1 / Fix audit setup deadlock by removing multiprocessing from middleware (#19979)
PR #19941 moved zettarepl out of the middleware process and removed the
setting that made multiprocessing start its workers as fresh processes.
That PR assumed zettarepl was the last user of multiprocessing in the
middleware process, but two users remained. Without the setting, their
workers became copies of the middleware process, and that allows a
deadlock when a pool of workers shuts down.
The deadlock works like this. When the pool shuts down, the thread that
owns the pool takes the lock that guards the queue of tasks and never
gives it back. It then waits for every worker to exit. A worker that is
still waiting for that lock can never get it, so the pool sends it a
termination signal. A fresh process dies from that signal. A copy of the
middleware process keeps the signal handling of the middleware, which
does nothing in a worker, so the worker stays alive. The thread waits
for the worker, the worker waits for the lock that the thread holds, and
neither can continue.
[23 lines not shown]
[AMDGPU] Make named barrier type 1 byte to fix barrier IDs of array elements
Since #209746, a pointer in the barrier address space (15) is the barrier ID,
and the backend reads the ID as `ptr & 0x3F`. But `target("amdgcn.named.barrier", 0)`
is still 16 bytes, so GEP to element `i` of a barrier array adds `16 * i` to the
barrier ID. For example, if `@bars` gets barrier ID 1, `&bars[2]` selects barrier
33 instead of 3.
This PR changes the type to 1 byte, so one array element is one barrier ID.
Fixes LCOMPILER-2898.
Fix audit setup deadlock by removing multiprocessing from middleware
PR #19941 moved zettarepl out of the middleware process and removed the
setting that made multiprocessing start its workers as fresh processes.
That PR assumed zettarepl was the last user of multiprocessing in the
middleware process, but two users remained. Without the setting, their
workers became copies of the middleware process, and that allows a
deadlock when a pool of workers shuts down.
The deadlock works like this. When the pool shuts down, the thread that
owns the pool takes the lock that guards the queue of tasks and never
gives it back. It then waits for every worker to exit. A worker that is
still waiting for that lock can never get it, so the pool sends it a
termination signal. A fresh process dies from that signal. A copy of the
middleware process keeps the signal handling of the middleware, which
does nothing in a worker, so the worker stays alive. The thread waits
for the worker, the worker waits for the lock that the thread holds, and
neither can continue.
[24 lines not shown]
[Scalar] Declare command line options in TableGen (#230370)
PipelineTuningOptions and LICMOptions read -forget-scev-loop-unroll,
-licm-mssa-optimization-cap, and -licm-mssa-max-acc-promotion through
getters instead of extern declarations.
Options read with getNumOccurrences() become std::optional members or
OptionalBoolField; readers that relied on the cl::init value apply it
with value_or()/valueOr(). -lsr-drop-solution becomes an
OptionalBoolField. CRCStrategyKind and MatrixLayoutTy move to
ScalarOptions.h.
Aided by Opus 5.5
[AMDGPU] Limit the fmul fusion discount to a matching context type
The fmul is free when its context instruction feeds a fusable fadd or
fsub. A vector fmul priced with a scalar lane as the context inherits
that fusion only when it is emitted lane by lane. On a packed type the
scalar user does not show that the vector fmul feeds a vector fadd, and
products extracted into a scalar fadd chain keep the packed fmul while
the fma is lost.
Open only the named pool when zpool checks a pool name
While tracing zpool get on one pool I noticed that it opens every
imported pool before the one named on the command line.
zpool get, zpool set and zpool iostat call is_pool() to tell a pool
name from a vdev name. It opened every imported pool to compare names,
so a command on a small pool also paid for every other pool on the
system. is_pool() now opens only the pool it is asked about.
On a system with three pools, one of them with 1200 disks, zpool get
all on the other two pools went from about 930 ms to 15 ms and 22 ms.
On the 1200 disk pool itself it stayed at about 1.8 seconds, because
the named pool is still opened twice, once for this check and once by
the command.
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Ameer Hamza <ameer.hamza at truenas.com>
Reviewed-by: Rob Norris <rob.norris at truenas.com>
Signed-off-by: Caleb St. John <yocalebo at gmail.com>
Closes #19271
[flang] Test portability warnings for negated literals and signed operands (#230514)
Two flang portability warnings had no test checking their text: "negated
maximum INTEGER(KIND=k) literal" (-Wbig-int-literals), emitted when
unary minus applied to an integer literal folds to the most negative
value of the kind, and "nonstandard usage: signed mult-operand", emitted
for the extension that accepts a sign before a mult-operand.
This adds the negated-literal cases to Semantics/int-literals.f90, next
to the signed-literal cases that must not warn, and a test_errors.py run
to Evaluate/signed-mult-opd.f90, which already checks the folded values
of the same expressions. Test-only change.
[SLP]Fix uitofp of signed demoted operand
uitofp reads its operand as unsigned, so a narrowed signed operand
must be sign-extended back first.
Fixes #230406
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/230583
[lldb][Fortran] Added support for base types to DWARFASTParserFortran, tests for DWARFASTParserFortran and a method to get the parser from TypeSystemFortran