MachinePipeliner: Pass instruction to findLoopIncrementValue (#219966)
The helper recovered the loop block from the operand's parent
instruction. Pass the containing instruction directly so it no longer depends
on MachineOperand::getParent().
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[flang][cuda] Reject DEVICE derived types with attributed allocatable components (#220411)
CUDA attributed allocatables are allocated from the host in most case
(PINNED, MANAGED, UNIFIED). A DEVICE derived-type object keeps its
component descriptors in device global memory, so allocating such a
component is not valid. It's also invalid for a DEVICE component unless
the allocate statement is in device context.
[VPlan] Allow non-live-in start values for widened inductions (NFC). (#220716)
Prepare VPWidenIntOrFpInductionRecipe for modeling the full epilogue
skeleton and resume values properly, by removing the VPIRValue
requirement for the start value.
The only requirement for the start value is that it dominates the phi,
which the verifier already ensures.
This is NFC today, but prepares for modeling the full epilogue skeleton
in VPlan, which requires adding phi nodes in the preheader before
execute.
[flang][OpenMP] Track reachable metadirective replacement paths
The existing semantic checks can validate loop-associated directives in a
METADIRECTIVE against the following loop, but they do not model how the
METADIRECTIVE chooses among its replacements.
Today each WHEN is considered independently: if its selector can match, its
replacement is checked. Selection instead ranks all applicable candidates as
a set. An unguarded higher-ranked candidate makes lower-ranked candidates
unreachable, while a dynamically guarded candidate leaves them reachable
when its condition is false. Treating both cases alike can diagnose loop
requirements on a replacement that can never be selected.
The selected replacement can also affect later selection. Its directive
contributes to the construct context seen by a nested metadirective. The
checker currently retains only syntactic nesting, so nested construct
selectors cannot observe a directive selected by an enclosing
metadirective.
[44 lines not shown]
[libc++][pstl] Implementation of parallel std::find_end() based on __parallel_find() (#218321)
This PR implements a parallel version of `std::find_end()` based on
`__parallel_find()`.
The algorithm crops the input range to a range where a potential match
can start and runs a chunked parallel find on the cropped range.
Inside each chunk potential matches are looked for using the serial
`std::find_end()` and the last one found is returned.
Since it's based on `__parallel_find()`, the algorithm supports early
termination.
Part of #99938.
[CIR] Use the modern enum case classes
The `I32EnumAttrCase` family carries an `Attr` half, and an `IntegerAttr`
predicate with it, that a CIR enum has no use for now that the enums derive
from `EnumInfo`. Upstream says of those forms that they "are not needed when
using the newer `EnumCase` form".
Rename all 198 of them to `I32EnumCase`, `I32BitEnumCaseNone`,
`I32BitEnumCaseBit` and `BitEnumCaseGroup`. The group class drops its width
prefix because the modern spelling takes the width from its cases.
NFC, mechanical.
[CIR] Record why the CUDA registration attribute parses itself
hasCustomAssemblyFormat with no explanation invites the question of whether a
declarative assemblyFormat would do. It would not. The three flags print as
presence-only keywords and parse in any order, while an optional group
anchored on a `bool` parameter parses and prints a value, so the group would
spell `extern true`. MLIR has no presence-only flag for `bool` in an
attribute format, unlike UnitAttr in an operation format. struct(params)
round-trips but spells the attribute
`<device_side_name = "i", kind = Variable, isExtern = true>` instead of
`<i, Variable, extern>`.
NFC.
[CIR] Drop the redundant suffix from the inline kind mnemonic
inline_kind was the one CIR enum attribute mnemonic still repeating what its
C++ enum class name says. The attribute now spells
`#cir.inline<always_inline>`. The operation argument keeps the name
inline_kind, since that is the accessor name, so the printed form reads
`inline_kind = #cir.inline<always_inline>`.
The enum's summary also becomes "inline kind" rather than the camelCase
"inlineKind", which is what generated docs show now that CIR_InlineKindAttr
no longer overrides it.
25 CHECK lines change across four test files. Nine are in an
aarch64-registered-target test, unsupported in an X86-only build, but the
substitution matches the two CIR tests that do run.
[CIR] Derive lowering attr names from cppClassName, not the def name
CIRLoweringEmitter built its CXX_ABI_ALWAYS_LEGAL_ATTRS entries with
GetOpCppClassName, which splits the TableGen def name at the first
underscore. That works only while every def is named CIR_<CppClassName>Attr.
When one is not, the emitter writes an `isa<>` for a class that does not
exist, and the failure lands as a compile error in generated code.
Attributes carry the authoritative name in cppClassName, which
GenerateAttrToValueVisitor was already reading. Factor that out as
GetAttrCppClassRef and use it for both attribute paths. GetOpCppClassName
stays for operations.
NFC, and checkable. No CIR attribute overrides cppClassName, so the generated
CIRLowering.inc is byte-identical.
[CIR] Drop dead ceremony around the CIR enum attributes
Five things that no longer earn their place in the CIR enum attribute
machinery.
CIR_CleanupKindAttr carried three. Its cppClassName restated the default
AttrDef already derives. Its skipDefaultBuilders plus hand-written
AttrBuilder existed only to default $value to CleanupKind::All, which no
caller relies on, so the generated builders stayed suppressed for nothing.
And its summary and description restated the name, overriding the enum's own
"cleanup kind" that EnumAttr would otherwise inherit. The isNormal, isEH and
isNormalAndEH helpers stay.
CIR_TLSModelAttr's summary restated its name the same way, so only that goes.
CIR_DefaultValuedEnumParameter has never had a user.
NFC.
[CIR] Move the CIR enums off the legacy EnumAttrInfo hierarchy
MLIR has two enum hierarchies. `EnumAttrInfo` doubles as an `IntegerAttr`
constraint, so every CIR enum had to clear `genSpecializedAttr` to say it did
not want one. `EnumInfo` describes a C++ enum and nothing more.
Derive the CIR bases from `I32Enum`, `I64Enum` and `I32BitEnum`, and widen
`CIR_EnumAttr` to the `EnumInfo` that upstream `EnumAttr` already takes. The
flag no longer exists to clear. `FPClassTestEnum` gets unquoted printing from
`BitEnumBase` rather than overriding `printBitEnumQuoted`, and
`CIR_KnownFuncKind` drops a `parameterPrinter` the generated `operator<<`
now covers, still spelling `#cir.func_identity<"std::find">`.
AMDGPU wraps an `I32Enum` in an `EnumAttr` with this same bracketed format.
Parsing moves to the generated `FieldParser`, whose diagnostic names the
accepted spellings, so two `expected-error` lines change. Generated attribute
code drops 16 KB as 28 inlined parsers collapse into it.
[CIR] Migrate the FPClassTest bit enum and unquote its flags
cir.is_fp_class printed its flags inconsistently. Single-bit values came out
bare, as in `fcSNan`, while group values and combinations came out quoted, as
in `"fcInf"` and `"fcSNan|fcNegInf"`. That comes from I32BitEnumAttr setting
printBitEnumQuoted, which EnumAttr.td keeps only for backwards compatibility.
Clearing the bit and using the `enum` directive selects the separator-aware
parser and printer, so every value now spells unquoted:
cir.is_fp_class %x, fcSNan|fcNegInf : (!cir.float) -> !cir.bool
The enum also drops its specialized IntegerAttr for a CIR_EnumAttr wrapper,
giving it the standalone spelling `#cir.fp_class<fcSNan|fcNegInf>`. This
changes operation syntax, so it updates 37 CHECK lines.
[CIR] Migrate GlobalLinkageKind, CallingConv and SideEffect off IntegerAttr
GlobalLinkageKind, CallingConv and SideEffect generated IntegerAttr
subclasses with no dialect spelling of their own. Each now sets
genSpecializedAttr = 0 and gains a CIR_EnumAttr wrapper, and cir.global wraps
$linkage in `enum()`. GlobalLinkageKind spells `#cir.linkage<internal>`,
dropping both the `global_` prefix and the `_kind` suffix.
cir.func and cir.call print all three by hand, but they stream
stringifyGlobalLinkageKind(getLinkage()) and friends, which take the enum
rather than the attribute, so those sites are unchanged.
Operation syntax is unchanged.
[CIR] Migrate AssumeBundleKind, AtomicFetchKind and AsmFlavor off IntegerAttr
AssumeBundleKind, AtomicFetchKind and AsmFlavor generated IntegerAttr
subclasses with no dialect spelling of their own. Each now sets
genSpecializedAttr = 0 and gains a CIR_EnumAttr wrapper.
Unlike the other CIR operation enums, these three are reached through
hand-written parsers and printers, so they needed checking individually.
cir.atomic.fetch references $binop declaratively and gains an `enum()`
wrapper. The other two need no change, since printAssumeBundle is already
typed on cir::AssumeBundleKindAttr and InlineAsmOp::print streams the enum
rather than the attribute.
Operation syntax is unchanged.
[CIR] Migrate MemOrder and SyncScopeKind off IntegerAttr
MemOrder and SyncScopeKind, the enums the atomic operations share, generated
IntegerAttr subclasses with no dialect spelling of their own.
Both now set genSpecializedAttr = 0 and gain CIR_EnumAttr wrappers, spelling
`#cir.mem_order<seq_cst>` and `#cir.sync_scope<system>`, and the atomic
operations wrap their arguments in `enum()` to keep the bare keyword.
`enum()` works as an optional-group anchor, so the `syncscope` and `atomic`
groups on cir.load and cir.store are unaffected. Operation syntax is
unchanged.
[CIR] Migrate seven operation enums off IntegerAttr
CastKind, DynamicCastKind, CmpOpKind, ComplexRangeKind, InitCatchKind,
CaseOpKind and AwaitKind generated IntegerAttr subclasses. The operations
printed them symbolically, but in an attribute dictionary `cir.cast bitcast`
was stored as `kind = 1 : i32`.
Each enum now sets genSpecializedAttr = 0 and gains a CIR_EnumAttr wrapper,
and the operations wrap the argument in `enum()` to keep the bare keyword,
giving spellings like `#cir.cast<bitcast>`. Mnemonics drop the suffix the C++
class name carries. DynamicCastKind spells out `dynamic_cast`, since
`dyn_cast` is taken by the operation and by `#cir.dyn_cast_info`.
CUDADeviceVarKind gets no wrapper, being only a raw parameter of
CIR_CUDAVarRegistrationInfoAttr.
Operation syntax is unchanged, and enum-attrs.cir covers the new spellings.
[CIR] Delete the unused cir::VisibilityAttr
CIR_VisibilityAttr had no users. `cir.global` and `cir.func` carry visibility
as `EnumProp<CIR_VisibilityKind>`, a property rather than an attribute, so
nothing ever built or printed the attribute.
Its only consumer was CIRGenModule::getGlobalVisibilityAttrFromDecl, itself
never called, and that was the only caller of
getGlobalVisibilityKindFromClangVisibility, so all three go together. The
similar getCIRVisibilityKind does have a caller and stays, as does
CIR_VisibilityKind, which the property is built from.
This also removes one of the two attributes overriding their assembly format
to a bare `$value`.
[CIR] Drop lang_address_space's custom parenthesized attribute format
CIR_LangAddressSpaceAttr overrode its assembly format to
`(` custom<AddressSpaceValue>($value) `)`. The parentheses defeated the
dialect's `#cir.mnemonic<...>` syntax, so the attribute printed as
`#cir<lang_address_space(offload_global)>`. It now uses the bracketed
CIR_EnumAttr default and spells `#cir.lang_address_space<offload_global>`.
The `lang_address_space(x)` spelling inside `!cir.ptr` and `cir.global` is
unaffected, since that goes through MemorySpaceAttrInterface in CIRTypes.cpp.
The attribute-level pair in CIRAttrs.cpp was only reachable from the deleted
format, so it goes away.
That left the CIRTypes.cpp pair sharing a name with the deleted enum hooks,
which reads as duplicated logic even though the two are unrelated: these take
a mlir::ptr::MemorySpaceAttrInterface and dispatch over LangAddressSpaceAttr
and TargetAddressSpaceAttr to spell the address space nested in a type or op,
so no ODS default could replace them. Rename them to
parseMemorySpace/printMemorySpace, matching MemorySpaceAttrInterface and
[4 lines not shown]
[CIR] Print cir.global's TLS model without naked angle brackets
`cir.global` printed `tls_model = <tls_dyn>`. Those brackets were the
leftover delimiters of `#cir.tls_model<tls_dyn>` after the printer stripped
the dialect prefix and mnemonic, so one enum had two spellings and the
per-global one was not something anyone would write by hand.
Wrapping the argument in the `enum` directive prints the symbolic value on
its own:
cir.global external tls_model = tls_dyn @a = #cir.int<5> : !s32i
The standalone attribute is unchanged. invalid-tls.cir now gets one
diagnostic from the enum parser instead of two.
[CIR] Give the cleanup kind a proper standalone attribute spelling
CleanupKindAttr overrode its assembly format to a bare `$value` so
`cir.cleanup.scope` would print `cleanup all`. The cost was that the
attribute had no readable standalone form, falling back to
`#cir<cleanup_kind all>`.
The `enum($attr)` operation directive removes the tradeoff. The attribute
keeps CIR_EnumAttr's bracketed default and now spells `#cir.cleanup<all>`,
while the operations ask for the bare keyword. The mnemonic drops the `_kind`
suffix the C++ class name carries.
Operation syntax is unchanged. invalid-loop-cleanup.cir now gets one
diagnostic from the enum parser instead of two.
[clang] Clean up switch stack when transforming an invalid body (#211162)
This PR resolves https://github.com/llvm/llvm-project/issues/210575.
Call RebuildSwitchStmt() irrespective of whether the body was succesfully
transformed to make sure we always pop an entry off the switch stack.
Assisted-by: OpenAI Codex
[InstCombine] Don't assume a trunc with nsw is never poison (#220930)
`impliesPoisonOrCond()` treats `trunc nuw X to i1` as never poison when
`X` is known to be in `[0, 1]`, which makes the assumed-poison premise
vacuously true and folds a logical and/or into a bitwise one.
`m_NUWTrunc()` matches on a flag subset, so it also matches `trunc nuw
nsw`. Such a trunc is still poison for `X == 1`, because `nsw` requires
the truncated bits to equal the top bit of the result, and they are zero
while the result is one. For `%c = false` and `%x = 1` the source
returns `false` while the folded form returns `poison`:
https://alive2.llvm.org/ce/z/Q6G5-v
None of the remaining `m_NUWTrunc()` users assumes that the matched
trunc is non-poison, so they are unaffected.
Fixes #220487.
[CIR][NFC] Share getSuccessorRegions across region-branch ops
Six of the ten CIR ops implementing RegionBranchOpInterface reported the same
successors: any of their regions may be entered from the parent operation, and
every region exit goes back to it. Add a CIR_EnterAnyRegionBranchOpBase class
that appends that definition to the one inherited from CIR_RegionBranchOpBase,
retarget the six ops onto it and delete their hand-written definitions.
The shared definition walks getRegions() rather than naming region accessors.
For all six ops the entry regions were exactly the declared regions in
declaration order, so it reports the same successors in the same order. The doc
comments of the deleted ScopeOp and TernaryOp definitions go with them, instead
of being left behind on the neighbouring builders.
IfOp, GlobalOp, TryOp and AwaitOp stay on the base class, since their successors
depend on the operation: IfOp falls back to the parent when the else region is
empty, GlobalOp skips its optional ctor and dtor regions, TryOp iterates
variadic handler regions, and AwaitOp routes ready to resume and suspend.
[CIR][NFC] Share getSuccessorInputs across region-branch ops
The ten CIR ops implementing RegionBranchOpInterface each hand-wrote
getSuccessorInputs, and all ten bodies were equivalent: regions take no
inputs, and returning to the parent yields the parent's results. Three did
not look equivalent but are: CleanupScopeOp and CoroBodyOp returned an empty
ValueRange unconditionally and declare no results, and AwaitOp returned
region block arguments but carries NoRegionArguments, so those ranges are
always empty.
Add a CIR_RegionBranchOpBase ODS class that declares the method and generates
the single shared body through extraClassDefinition, mirroring the existing
CIR_LoopOpBase, and retarget all ten ops onto it.
The generated CIROps.h.inc is unchanged and CIROps.cpp.inc gains exactly the
ten definitions removed from CIRDialect.cpp.
[CIR] Add RegionBranchOpInterface unit tests and fix cir.await successors
Five of the ten ops implementing RegionBranchOpInterface have no unit test
coverage: cir.case, cir.cleanup.scope, cir.global, cir.await and
cir.coro.body. Add tests for all five.
Covering cir.await exposes a disagreement with its own terminator.
cir.condition terminates the ready region and reports {resume, suspend} as
its successors when the parent is an await, but AwaitOp::getSuccessorRegions
listed all three regions as entry successors and reported the parent op as
the successor of every region exit. Fix it to match cir.condition: ready is
the only entry successor, exiting ready branches to resume or suspend, and
exiting suspend or resume returns to the parent operation.
cir.await declares no results and carries NoRegionArguments, so successor
operand and input counts stay at zero along every edge and the MLIR verifier
is unaffected.
[llvm-exegesis] Use %t.o instead of a literal %d in a RISC-V test (#220971)
The name therefore collides as soon as a second test in that directory
uses it. llvm-objdump mmaps the file while the concurrently running test
truncates and rewrites it, and touching a page past the new end of the
mapping kills llvm-objdump with SIGBUS, with no diagnostic.
Use %t.o, which lit expands to a name unique to the test, as every other
llvm-exegesis test already does.
[X86] Prefer AVX512 VPCMP against zero over splat(1) for sle/slt (#216716)
InstCombine canonicalizes `icmp sle x, 0` to `icmp slt x, 1`. AVX512 `VPCMP`
can encode LE/GE against a zeroed register, but we were loading splat(1)
from the constant pool (`vpcmpltb .LCPI`).
### Approach
In `combineSetCC`, for AVX512 `vXi1` integer compares:
- Rewrite `slt x, splat(1)` / `sgt splat(1), x` to `sle x, 0`
- Rewrite `sgt x, splat(-1)` / `slt splat(-1), x` to `sge x, 0`
- Skip the existing LE/GE → LT/GT `incDecVectorConstant` fold when it
would
replace a zero splat with ±1 (avoids oscillating with the rewrite above)
This is the vector analog of scalar `TranslateX86CC` (`SETLT x, 1` →
`COND_LE`
vs 0). InstCombine is left unchanged.
Pre-AVX512 (`-mcpu=x86-64-v3`) still uses `vpcmpgtb` vs zero + `not`.
[9 lines not shown]
[libc] Make File::FileLock public in File data structure (#221005)
Commit 9db0037bf1b3 moved FileLock from namespace scope into class File
so internal methods could use RAII locking, but placed it in the private
section.
Move FileLock to the public section of class File so that external
callers operating on File streams can use RAII locking instead of manual
lock and unlock calls.
Port callers in puts, fgets, perror, fgetws, reopenfile, and the pwd
flat file database parser to use File::FileLock, and add a unit test
verifying RAII locking on File streams.
Assisted-by: Automated tooling, human reviewed.