LLVM/project 5a1dbdd — compiler-rt/lib/builtins/hexagon ieee_qfloat_fast_convert_hf_to_ub_rne.S ieee_qfloat_fast_convert_hf_to_b_rne.S

[Hexagon] Add v79 QFloat HVX runtime conversion funcs (#229481)

Add four assembly files for half-float to integer conversions using
QFloat (qf16) on Hexagon V79.

Each function provides IEEE round-to-nearest-even or fast
round-half-away-from-zero conversion for signed and unsigned 8-bit and
16-bit integer destinations.

V79 lacks direct half-float to integer conversion. V75 and below have
native IEEE vcvt instructions,
while V81 and newer have dedicated hardware support.


Co-authored-by: Sumanth Gundapaneni <sgundapa at quicinc.com>
DeltaFile
+154-0compiler-rt/lib/builtins/hexagon/ieee_qfloat_conv_hf_to_uh_rne.S
+133-0compiler-rt/lib/builtins/hexagon/ieee_qfloat_conv_hf_to_b_rne.S
+133-0compiler-rt/lib/builtins/hexagon/ieee_qfloat_conv_hf_to_ub_rne.S
+78-0compiler-rt/lib/builtins/hexagon/ieee_qfloat_conv_hf_to_h_rne.S
+63-0compiler-rt/lib/builtins/hexagon/ieee_qfloat_fast_convert_hf_to_b_rne.S
+61-0compiler-rt/lib/builtins/hexagon/ieee_qfloat_fast_convert_hf_to_ub_rne.S
+622-03 files not shown
+718-09 files

LLVM/project bb26a17 — llvm/lib/Transforms/InstCombine InstCombineCasts.cpp, llvm/test/Transforms/InstCombine fptrunc-bfloat-to-half.ll

[InstCombine] Don't narrow FP ops when the destination range is too small (#230379)

Fixes #229651.

InstCombine narrows `fptrunc (binop (fpext X), (fpext Y))` into a binop
in the destination type, but it only compared the significand widths.
`bfloat` has fewer significand bits than `half` but a much larger
exponent range, so `bfloat` operands were narrowed to `half` and large
values overflowed to infinity:

```llvm
%x = fpext bfloat %a to double
%y = fpext bfloat %b to double
%sum = fadd double %x, %y
%r = fptrunc double %sum to half
```

For `65536 + -65536` this returns `+0.0`, but after narrowing both
operands become `inf` and `-inf`, and the sum is NaN.

    [4 lines not shown]
DeltaFile
+142-0llvm/test/Transforms/InstCombine/fptrunc-bfloat-to-half.ll
+15-3llvm/lib/Transforms/InstCombine/InstCombineCasts.cpp
+157-32 files

LLVM/project dff40f3 — llvm/lib/Target/AMDGPU AMDGPUCoExecInfo.h

[AMDGPU] NFC: Drop constexpr from getFlavorName and getCoExecMask (#230334)

Both functions end in llvm_unreachable. GCC 8 rejects a constexpr
function that can reach a call to llvm_unreachable_internal when
assertions are enabled, which breaks the clang-ppc64le-linux-test-suite
and clang-ppc64le-linux-multistage builders.

#203603 already fixed this for getFlavorName, but the constexpr came
back when the function moved into AMDGPUCoExecInfo.h in #204077. Neither
function is used in a constant expression.
DeltaFile
+2-2llvm/lib/Target/AMDGPU/AMDGPUCoExecInfo.h
+2-21 files

LLVM/project 8e242e6 — clang/test/CodeGen/AArch64 abi-classify-pure-scalable.c, llvm/include/llvm/ABI FunctionInfo.h

[LLVMABI][AARCH64] Implement Pure Scalable Type handling (#227504)

This change adds support for classifying Pure Scalable Type arguments
for AArch64 targets in the LLVM ABI library. Pure Scalable Types are
passed in registers, expanded if necessary, unless the argument is
unnamed or there are not sufficient registers available, in which case
they are passed indirectly. Pure Scalable Types are treated as single
named arguments when used as return types.

Assisted-by: Cursor / various models
DeltaFile
+416-0llvm/unittests/ABI/AArch64TargetInfoTest.cpp
+222-0llvm/unittests/ABI/TargetInfoTest.cpp
+210-0llvm/lib/ABI/Targets/AArch64.cpp
+204-0clang/test/CodeGen/AArch64/abi-classify-pure-scalable.c
+149-0llvm/lib/ABI/TargetInfo.cpp
+43-1llvm/include/llvm/ABI/FunctionInfo.h
+1,244-15 files not shown
+1,321-111 files

LLVM/project 8b087a7 — llvm/lib/Transforms/HipStdPar HipStdPar.cpp, llvm/test/Transforms/HipStdPar math-fixup.ll

HipStdPar: Stop redirecting f64 exp and exp2 (#230573)

The AMDGPU backend now expands llvm.exp.f64 and llvm.exp2.f64 directly.
DeltaFile
+2-2llvm/test/Transforms/HipStdPar/math-fixup.ll
+0-2llvm/lib/Transforms/HipStdPar/HipStdPar.cpp
+2-42 files

LLVM/project 0258719 — llvm/lib/Target/Hexagon HexagonGlobalScheduler.cpp

[Hexagon][NFC] Fix duplicate A2_tfr check in HexagonGlobalScheduler (#227762)

Fix a typo where both sides of a logical AND checked the same opcode
(A2_tfr). The second check should be A2_tfrsi.
DeltaFile
+1-2llvm/lib/Target/Hexagon/HexagonGlobalScheduler.cpp
+1-21 files

LLVM/project baf56e3 — llvm/lib/Target/RISCV RISCVFrameLowering.cpp, llvm/test/CodeGen/RISCV stack-clash-prologue.ll

[RISCV] Correct the CFA offsets for the stack probe loop. (#230322)

We need to take into account that we may have already done a
FirstSPAdjust. This is the same bug as #164805, which was fixed
for the unrolled probes in #166616 but not for the probe loop.

Fixes #230291.

Co-Authored-By: Claude Opus 5.5 <noreply at anthropic.com>
DeltaFile
+69-0llvm/test/CodeGen/RISCV/stack-clash-prologue.ll
+5-4llvm/lib/Target/RISCV/RISCVFrameLowering.cpp
+74-42 files

LLVM/project 4d5231f — utils/bazel/llvm-project-overlay/llvm BUILD.bazel

[Bazel][llvm] Generate InlinerModels and RegAllocEvictModels headers (#230576)

Fixes build failures following commit c90b6414e607 ([mlgo] Allow passing
pre-emitc-ed models (#227941)) where InlinerModels.h and
RegAllocEvictModels.h are unconditionally included in MLInlineAdvisor
and MLRegAllocEvictAdvisor.

LLM-aided
DeltaFile
+84-3utils/bazel/llvm-project-overlay/llvm/BUILD.bazel
+84-31 files

LLVM/project 254efe1 — llvm/lib/Target/AArch64 AArch64ISelLowering.cpp

[AArch64] Add NVCAST to inputs of AArch64ISD::REV32/REV64 to match SDTypeProfile. (#230211)

The SDTypeProfile says the input type should match the result type. Add
NVCASTs to satisfy this constraint.

Found by adding SDTCisSameAs to SDNodeInfo::verifyNode.
DeltaFile
+15-5llvm/lib/Target/AArch64/AArch64ISelLowering.cpp
+15-51 files

LLVM/project 9a9abb8 — mlir/lib/Dialect/X86/Transforms VectorContractToAMXDotProduct.cpp, mlir/test/Dialect/X86/AMX vector-contract-to-tiled-dp.mlir

[mlir][x86] Fix VNNI operand load offsets (#230205)

Fixes the tile load offsets of VNNI operands in the AMX contraction
lowering.

An offset into the VNNI packed dims was not scaled by the VNNI factor,
so loads at a non-zero offset read the wrong data. VNNI operands whose
buffer's innermost dim is not the static VNNI factor are now rejected,
as their tiles cannot be loaded directly.

Assisted-by: Claude
DeltaFile
+427-0mlir/test/Dialect/X86/AMX/vector-contract-to-tiled-dp.mlir
+52-17mlir/lib/Dialect/X86/Transforms/VectorContractToAMXDotProduct.cpp
+479-172 files

FreeNAS/freenas 6f74a62 — src/middlewared/middlewared/plugins/webshare sharing.py

Store dataset and relative_path when creating a WebShare share
DeltaFile
+3-0src/middlewared/middlewared/plugins/webshare/sharing.py
+3-01 files

LLVM/project 80e080a — libcxx/test/std/language.support/cmp/cmp.type type_order.compile.pass.cpp

[libc++] Mark type_order test as unsupported on recent AppleClang (#230210)

The most recent AppleClang versions don't support that feature yet.
DeltaFile
+1-1libcxx/test/std/language.support/cmp/cmp.type/type_order.compile.pass.cpp
+1-11 files

FreeNAS/freenas d7badbb — src/middlewared/middlewared/plugins/audit utils.py, src/middlewared/middlewared/plugins/disk_ retaste.py

NAS-144385 / 28.0.0-BETA.1 / Fix audit setup deadlock by removing multiprocessing from middleware (#19979)

PR #19941 moved zettarepl out of the middleware process and removed the
setting that made multiprocessing start its workers as fresh processes.
That PR assumed zettarepl was the last user of multiprocessing in the
middleware process, but two users remained. Without the setting, their
workers became copies of the middleware process, and that allows a
deadlock when a pool of workers shuts down.

The deadlock works like this. When the pool shuts down, the thread that
owns the pool takes the lock that guards the queue of tasks and never
gives it back. It then waits for every worker to exit. A worker that is
still waiting for that lock can never get it, so the pool sends it a
termination signal. A fresh process dies from that signal. A copy of the
middleware process keeps the signal handling of the middleware, which
does nothing in a worker, so the worker stays alive. The thread waits
for the worker, the worker waits for the lock that the thread holds, and
neither can continue.


    [23 lines not shown]
DeltaFile
+15-25src/middlewared/middlewared/plugins/disk_/retaste.py
+7-1src/middlewared/middlewared/plugins/audit/utils.py
+22-262 files

LLVM/project edb27a6 — clang/test/CodeGenHIP amdgpu-barrier-type.hip, clang/test/SemaCXX amdgpu-barrier.cpp

[AMDGPU] Make named barrier type 1 byte to fix barrier IDs of array elements

Since #209746, a pointer in the barrier address space (15) is the barrier ID,
and the backend reads the ID as `ptr & 0x3F`. But `target("amdgcn.named.barrier", 0)`
is still 16 bytes, so GEP to element `i` of a barrier array adds `16 * i` to the
barrier ID. For example, if `@bars` gets barrier ID 1, `&bars[2]` selects barrier
33 instead of 3.

This PR changes the type to 1 byte, so one array element is one barrier ID.

Fixes LCOMPILER-2898.
DeltaFile
+10-78llvm/test/CodeGen/AMDGPU/s-barrier-signal-var-gep.ll
+14-22llvm/test/CodeGen/AMDGPU/s-barrier-array-index.ll
+5-5clang/test/CodeGenHIP/amdgpu-barrier-type.hip
+2-4llvm/lib/IR/Type.cpp
+2-2clang/test/SemaHIP/amdgpu-barrier.hip
+2-2clang/test/SemaCXX/amdgpu-barrier.cpp
+35-1134 files not shown
+39-11710 files

LLVM/project 3b00a39 — llvm/test/CodeGen/AMDGPU s-barrier-array-index.ll

[NFC][AMDGPU] Add test for indexing into named barrier arrays (#230515)
DeltaFile
+152-0llvm/test/CodeGen/AMDGPU/s-barrier-array-index.ll
+152-01 files

FreeNAS/freenas d5dd011 — src/middlewared/middlewared/plugins/audit utils.py, src/middlewared/middlewared/plugins/disk_ retaste.py

Fix audit setup deadlock by removing multiprocessing from middleware

PR #19941 moved zettarepl out of the middleware process and removed the
setting that made multiprocessing start its workers as fresh processes.
That PR assumed zettarepl was the last user of multiprocessing in the
middleware process, but two users remained. Without the setting, their
workers became copies of the middleware process, and that allows a
deadlock when a pool of workers shuts down.

The deadlock works like this. When the pool shuts down, the thread that
owns the pool takes the lock that guards the queue of tasks and never
gives it back. It then waits for every worker to exit. A worker that is
still waiting for that lock can never get it, so the pool sends it a
termination signal. A fresh process dies from that signal. A copy of the
middleware process keeps the signal handling of the middleware, which
does nothing in a worker, so the worker stays alive. The thread waits
for the worker, the worker waits for the lock that the thread holds, and
neither can continue.


    [24 lines not shown]
DeltaFile
+15-25src/middlewared/middlewared/plugins/disk_/retaste.py
+7-1src/middlewared/middlewared/plugins/audit/utils.py
+22-262 files

LLVM/project e61035a — llvm/lib/Transforms/Scalar DeadStoreElimination.cpp InductiveRangeCheckElimination.cpp

[Scalar] Declare command line options in TableGen (#230370)

PipelineTuningOptions and LICMOptions read -forget-scev-loop-unroll,
-licm-mssa-optimization-cap, and -licm-mssa-max-acc-promotion through
getters instead of extern declarations.

Options read with getNumOccurrences() become std::optional members or
OptionalBoolField; readers that relied on the cl::init value apply it
with value_or()/valueOr(). -lsr-drop-solution becomes an
OptionalBoolField. CRCStrategyKind and MatrixLayoutTy move to
ScalarOptions.h.

Aided by Opus 5.5
DeltaFile
+582-0llvm/lib/Transforms/Scalar/ScalarOptions.td
+84-135llvm/lib/Transforms/Scalar/LoopStrengthReduce.cpp
+79-125llvm/lib/Transforms/Scalar/SimpleLoopUnswitch.cpp
+51-151llvm/lib/Transforms/Scalar/LoopUnrollPass.cpp
+71-106llvm/lib/Transforms/Scalar/InductiveRangeCheckElimination.cpp
+38-100llvm/lib/Transforms/Scalar/DeadStoreElimination.cpp
+905-61752 files not shown
+1,540-1,79558 files

LLVM/project 3cc3220 — llvm/lib/Target/AMDGPU AMDGPUTargetTransformInfo.cpp, llvm/test/Transforms/SLPVectorizer/AMDGPU alt-fmul-fadd-cost.ll fma-operand-contract-selection.ll

[AMDGPU] Limit the fmul fusion discount to a matching context type

The fmul is free when its context instruction feeds a fusable fadd or
fsub. A vector fmul priced with a scalar lane as the context inherits
that fusion only when it is emitted lane by lane. On a packed type the
scalar user does not show that the vector fmul feeds a vector fadd, and
products extracted into a scalar fadd chain keep the packed fmul while
the fma is lost.
DeltaFile
+88-0llvm/test/Transforms/SLPVectorizer/AMDGPU/fmul-extract-fadd-chain.ll
+64-16llvm/test/Transforms/SLPVectorizer/AMDGPU/ordered-reduction-coalesced-loads.ll
+38-37llvm/test/Transforms/SLPVectorizer/AMDGPU/ordered-reduction-fma-fusion.ll
+36-6llvm/test/Transforms/SLPVectorizer/AMDGPU/fma-operand-contract-selection.ll
+7-11llvm/test/Transforms/SLPVectorizer/AMDGPU/alt-fmul-fadd-cost.ll
+5-2llvm/lib/Target/AMDGPU/AMDGPUTargetTransformInfo.cpp
+238-721 files not shown
+239-737 files

LLVM/project 3c5889b — llvm/test/Transforms/SLPVectorizer/AMDGPU ordered-reduction-coalesced-loads.ll fmul-extract-fadd-chain.ll

update tests
DeltaFile
+0-88llvm/test/Transforms/SLPVectorizer/AMDGPU/fmul-extract-fadd-chain.ll
+16-64llvm/test/Transforms/SLPVectorizer/AMDGPU/ordered-reduction-coalesced-loads.ll
+16-1522 files

OpenZFS/src e61fe25 — cmd/zpool zpool_main.c

Open only the named pool when zpool checks a pool name

While tracing zpool get on one pool I noticed that it opens every
imported pool before the one named on the command line.

zpool get, zpool set and zpool iostat call is_pool() to tell a pool
name from a vdev name. It opened every imported pool to compare names,
so a command on a small pool also paid for every other pool on the
system. is_pool() now opens only the pool it is asked about.

On a system with three pools, one of them with 1200 disks, zpool get
all on the other two pools went from about 930 ms to 15 ms and 22 ms.
On the 1200 disk pool itself it stayed at about 1.8 seconds, because
the named pool is still opened twice, once for this check and once by
the command.

Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Reviewed-by: Ameer Hamza <ameer.hamza at truenas.com>
Reviewed-by: Rob Norris <rob.norris at truenas.com>
Signed-off-by: Caleb St. John <yocalebo at gmail.com>
Closes #19271
DeltaFile
+14-12cmd/zpool/zpool_main.c
+14-121 files

LLVM/project 4ef9fa0 — flang/test/Evaluate signed-mult-opd.f90, flang/test/Semantics int-literals.f90

[flang] Test portability warnings for negated literals and signed operands (#230514)

Two flang portability warnings had no test checking their text: "negated
maximum INTEGER(KIND=k) literal" (-Wbig-int-literals), emitted when
unary minus applied to an integer literal folds to the most negative
value of the kind, and "nonstandard usage: signed mult-operand", emitted
for the extension that accepts a sign before a mult-operand.

This adds the negated-literal cases to Semantics/int-literals.f90, next
to the signed-literal cases that must not warn, and a test_errors.py run
to Evaluate/signed-mult-opd.f90, which already checks the folded values
of the same expressions. Test-only change.
DeltaFile
+12-0flang/test/Semantics/int-literals.f90
+5-0flang/test/Evaluate/signed-mult-opd.f90
+17-02 files

LLVM/project 65963d3 — llvm/lib/Transforms/Vectorize SLPVectorizer.cpp, llvm/test/Transforms/SLPVectorizer/X86 minbw-uitofp-signed-operand.ll

[SLP]Fix uitofp of signed demoted operand

uitofp reads its operand as unsigned, so a narrowed signed operand
must be sign-extended back first.

Fixes #230406

Reviewers: 

Pull Request: https://github.com/llvm/llvm-project/pull/230583
DeltaFile
+13-0llvm/lib/Transforms/Vectorize/SLPVectorizer.cpp
+4-2llvm/test/Transforms/SLPVectorizer/X86/minbw-uitofp-signed-operand.ll
+17-22 files

LLVM/project cd70216 — lldb/source/Plugins/SymbolFile/DWARF DWARFASTParserFortran.cpp, lldb/unittests/SymbolFile/DWARF DWARFASTParserFortranTests.cpp

[lldb][Fortran] Moved break in case and fixed failing test
DeltaFile
+7-3lldb/source/Plugins/SymbolFile/DWARF/DWARFASTParserFortran.cpp
+4-4lldb/unittests/SymbolFile/DWARF/DWARFASTParserFortranTests.cpp
+11-72 files

LLVM/project 8aee0da — lldb/source/Plugins/SymbolFile/DWARF DWARFASTParserFortran.cpp, lldb/unittests/SymbolFile/DWARF DWARFASTParserFortranTests.cpp

[lldb][Fortran] Addressed feedback, including early exits throughout ParseTypeFromDWARF and not treating DW_ATE_unsigned as a sentinel value
DeltaFile
+86-77lldb/source/Plugins/SymbolFile/DWARF/DWARFASTParserFortran.cpp
+4-4lldb/unittests/SymbolFile/DWARF/DWARFASTParserFortranTests.cpp
+90-812 files

LLVM/project 033594a — lldb/source/Plugins/SymbolFile/DWARF DWARFASTParserFortran.h DWARFASTParserFortran.cpp, lldb/source/Plugins/TypeSystem/Fortran TypeSystemFortran.h TypeSystemFortran.cpp

[lldb][Fortran] Added support for base types to DWARFASTParserFortran, tests for DWARFASTParserFortran and a method to get the parser from TypeSystemFortran
DeltaFile
+209-0lldb/unittests/SymbolFile/DWARF/DWARFASTParserFortranTests.cpp
+127-4lldb/source/Plugins/SymbolFile/DWARF/DWARFASTParserFortran.cpp
+17-1lldb/source/Plugins/SymbolFile/DWARF/DWARFASTParserFortran.h
+8-0lldb/source/Plugins/TypeSystem/Fortran/TypeSystemFortran.cpp
+3-0lldb/source/Plugins/TypeSystem/Fortran/TypeSystemFortran.h
+2-0lldb/unittests/SymbolFile/DWARF/CMakeLists.txt
+366-56 files

LLVM/project f1a1aa4 — lldb/source/Plugins/TypeSystem/Fortran TypeSystemFortran.cpp

Removed GetTypeName switch default
DeltaFile
+2-1lldb/source/Plugins/TypeSystem/Fortran/TypeSystemFortran.cpp
+2-11 files

LLVM/project fa3d2f0 — lldb/source/Plugins/TypeSystem/Fortran TypeSystemFortran.h TypeSystemFortran.cpp, lldb/unittests/Symbol TestTypeSystemFortran.cpp

[lldb][Fortran] Added spacing after 1-line ifs, inlined type cases and added default name for all base types
DeltaFile
+33-25lldb/source/Plugins/TypeSystem/Fortran/TypeSystemFortran.cpp
+2-2lldb/source/Plugins/TypeSystem/Fortran/TypeSystemFortran.h
+3-0lldb/unittests/Symbol/TestTypeSystemFortran.cpp
+38-273 files

LLVM/project d442642 — lldb/source/Plugins/TypeSystem/Fortran TypeSystemFortran.cpp, lldb/unittests/Symbol TestTypeSystemFortran.cpp

[lldb][Fortran] Removed eBasicTypeLongDoubleComplex handling, changed argument to snake_case and changes type creation logic
DeltaFile
+12-10lldb/source/Plugins/TypeSystem/Fortran/TypeSystemFortran.cpp
+0-2lldb/unittests/Symbol/TestTypeSystemFortran.cpp
+12-122 files

LLVM/project 2131fd9 — lldb/source/Plugins/TypeSystem/Fortran FortranTypes.cpp TypeSystemFortran.h, lldb/unittests/Symbol CMakeLists.txt TestTypeSystemFortran.cpp

[lldb][Fortran] Added base type support to TypeSystemFortran and Tests for TypeSystemFortran
DeltaFile
+237-0lldb/unittests/Symbol/TestTypeSystemFortran.cpp
+207-0lldb/source/Plugins/TypeSystem/Fortran/TypeSystemFortran.cpp
+76-0lldb/source/Plugins/TypeSystem/Fortran/FortranTypes.h
+32-35lldb/source/Plugins/TypeSystem/Fortran/TypeSystemFortran.h
+16-0lldb/source/Plugins/TypeSystem/Fortran/FortranTypes.cpp
+2-0lldb/unittests/Symbol/CMakeLists.txt
+570-351 files not shown
+571-357 files

LLVM/project 4558c7a — lldb/source/Plugins/TypeSystem/Fortran FortranTypes.h TypeSystemFortran.cpp, lldb/unittests/Symbol TestTypeSystemFortran.cpp

[lldb][Fortran] Added handling for the unsigned type and included complex types in IsFloatingPointType checks
DeltaFile
+54-1lldb/unittests/Symbol/TestTypeSystemFortran.cpp
+33-5lldb/source/Plugins/TypeSystem/Fortran/TypeSystemFortran.cpp
+1-0lldb/source/Plugins/TypeSystem/Fortran/FortranTypes.h
+88-63 files