[AMDGPU][GlobalISel] Legalize odd bf16 vectors (#226707)
Simplify `moreElementsIf` and `clampMaxNumElements` to just
`clampMaxNumElementsStrict` for `MinNumMaxNumIeee` and `MinNumMaxNum`.
Use `clampMaxNumElementsStrict` for `FSubActions`. This is the only
functional change.
Add testing for v3bf16 and v5b16.
---------
Signed-off-by: John Lu <John.Lu at amd.com>
[mlir] Avoid rewriting unreachable blocks in the greedy driver
A rewrite can disconnect a block after the iteration's initial CFG
sweep. Queued operations can develop self-referential SSA uses in
unreachable code, causing crashes or repeated rewrites that prevent the
worklist pass from returning. Track reachability through rewriter
notifications and skip these operations until the next iteration removes
their blocks.
The reproducer in #221152 exposes this gap in the initial sweep added by
#153957 for #153732. #154038 similarly skips unreachable blocks in the
walk-based driver. Reachable graph-region self-cycles in #194824 and
#205064 remain separate; #207185 proposes a fold-specific fix for them.
A forwarding ReachabilityListener sits between the rewriter and the
worklist driver. Notifications only record what changed: terminator
changes and removed blocks are recorded per block. Queries happen
between rewrites, once per popped operation, and bring the cache up to
date. A region without cached state, or with a different entry block
[38 lines not shown]
[RISCV] Remove riscv_mulh_i32/riscv_mulhu_u32 intrinsics. (#227819)
These are redundant with the llvm.smulh/umulh intrinsics that were
recently added.
This changes codegen because smulh/umulh are currently generically type
legalized with extends+mul+srli instead of using pmulh(u).w. This isn't
always a regression. Sometimes we are able to prove the inputs are
already extended or we use mul(u).w00 and sometimes we needed a sext.w
after the pmulh(u).w. I will handle this as a follow up as it will
affect more tests.
cleanup: Remove expired cfengine326 ports:
2026-09-30 sysutils/cfengine-masterfiles326: FreeBSD will only support N and N-1
2026-09-30 sysutils/cfengine326: FreeBSD will only support N and N-1
[bazel] Add `USE_CLANG_CL` for proper rules_cc toolchain resolution (#227859)
This change updates the `rules_cc` auto-configured toolchain to more
accurately represent a clang-cl toolchain.
Relates to: https://github.com/bazelbuild/rules_cc/pull/858
[SLP]Ignore the extracts in the instruction-count with high register pressure
In loops, the instruction-count check scales the tree entries by the
trip count, but not the extracts for the external users, so the extracts
just break the ties of the scaled counts.
With the fused fmul/fadd costing such ties reject the relaxation trees
of the reported kernel, the loop body stays scalar and spills. If more
values of the tree type are live in the block than there are registers,
ignore the extracts when they are the only reason to reject the tree.
Fixes the regression, reported in
https://github.com/llvm/llvm-project/pull/226117#issuecomment-5913566646
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/227864
cleanup: Remove expired cfengine324 ports:
2026-09-30 sysutils/cfengine-masterfiles324: FreeBSD will only support N and N-1
2026-09-30 sysutils/cfengine324: FreeBSD will only support N and N-1
[RISC-V][RVY] Rank 'y' after 'i' and 'e' in canonical extension order
Because 'y' is only valid as a base ISA letter in parseArchString and
not in AllStdExts, singleLetterExtensionRank previously fell through to
the unknown single-letter extension case and placed 'y' (and 'zy*'
extensions) after all standard single-letter extensions. Rank 'y'
immediately after 'i' and 'e' so that normalized ISA strings place 'y'
before 'm', 'a', 'f', 'd', 'c' and 'zy*' right after 'zi*'.
This commit was created with the help of AI tools
cxgbe: Add a sysctl/tunable to control KTLS offload of AES-CBC cipher suites
Disable these by default as they are less efficient and rarely used.
Sponsored by: Chelsio Communications
cleanup: Remove expired cfengine325 ports:
2026-09-30 sysutils/cfengine-masterfiles325: FreeBSD will only support N and N-1
2026-09-30 sysutils/cfengine325: FreeBSD will only support N and N-1
[AMDGPU] Price scalar integer to fp casts by source width and sign
Scalar sources between a byte and 31 bits fell to the default cost of one
while the matching vector lanes were already priced, which skewed the
difference SLP weighs a bundle against. Such a source is extended before
the conversion, and what the extension takes depends on the width, on the
sign and on whether the subtarget has SDWA and 16 bit instructions.
Sources narrower than a byte are left alone, because their vector form is
not priced either.
[AMDGPU] Price narrow integer to bfloat vector casts (#225342)
A vector lane of 9 to 15 or 17 to 31 bits converted to bfloat got the generic cost, which leaves out the rounding. Such a lane is converted to f32 first like any other narrow lane, so price it as the f32 conversion of the same source plus the rounding. Lanes of 8 and 16 bits keep their cost.
[RISC-V][MC] Reject x0 as the address/temporary register for load/store pseudos (#227512)
Using the zero register as the destination for integer load pseudos or as
the temporary register for floating-point load and store pseudos is
illegal since it is used to synthesize the target address, and using x0
would mean the auipc/qc.e.li result is ignored and we access an invalid
location.
This commit was created with the help of AI tools
Pull-Request: https://github.com/llvm/llvm-project/pull/227512