[Sparc][NFC] test sparc variadic aggregate handling (#216521)
equivalent of https://github.com/llvm/llvm-project/pull/216509 for
sparc. It similarly has some bugs passing aggregates with floats.
[Clang] Accept auto casts pre-C++23 as an extension (#200675)
GCC already supports this as an extension pre-C++23. It is also useful
for libc++ to replace `_LIBCPP_AUTO_CAST`.
Fixes #115609
[libc++] Simplify the implementation of std::make_from_tuple (#215067)
This does two major things:
1) It removes conditionals for C++20/pre-C++20. I don't understand why
this has ever been done. This made the implementation significantly more
complicated without any indication that it actually improved anything.
2) std::apply is used for expanding the tuple
[lldb] Fix invalid UTF-8 in JSON log message (#216185)
Enabling the JSON packet log part way through a session aborts an
assertions build:
```
(lldb) b f
(lldb) run
(lldb) log enable -j -f /tmp/pk.json gdb-remote packets
(lldb) next
Assertion failed: (false && "Invalid UTF-8 in value used as JSON"), function Value, file JSON.h, line 333.
```
`Log::EmitJSONMessage` passed the message straight to
`llvm::json::Value`,
which asserts on ill-formed UTF-8 and only then falls back to `fixUTF8`.
So a release build repairs the message while an assertions build dies.
The bytes come from the saved packets. Once logging is turned on,
[10 lines not shown]
Revert "[Clang] Support friend declarations with a dependent nested-name-specifier" (#216549)
Reverts llvm/llvm-project#208345
---
Revert dependent friend support due to GCC build failure
[VectorCombine] Drop invariant.group from scalarized stores (#212473)
!invariant.group is tied to a pointer SSA value, so it cannot be copied
from a vector store to a scalar store that uses a newly created GEP.
Drop the metadata after copying the remaining store metadata and update
the regression expectations.
Fixes https://github.com/llvm/llvm-project/issues/212472
Assisted-by: Codex
[clang][bytecode] Fix assertion failure in in valid continue stmt (#216547)
We need to handle the missing TargetLabel here, similarly to what we do
in break statements.
[MLIR][LLVM] Preserve pointer-valued metadata operands on import (#215743)
convertMetadataToAttrImpl only modelled ConstantInt operands wrapped in
a ConstantAsMetadata, so any metadata node containing a pointer constant
could not be represented and the whole node was rejected.
Add #llvm.md_null and #llvm.md_addrspacecast to model
ConstantPointerNull and addrspacecast constant expressions, keeping the
address space so that `ptr null` and `ptr addrspace(1) null` stay
distinct. MDAddrSpaceCastAttr verifies that its operand is itself
pointer-valued metadata.
Global values are constants, so ValueAsMetadata::get wraps them in a
ConstantAsMetadata and they never reached the ValueAsMetadata case.
Match them in the ConstantAsMetadata case instead, which also
generalizes the existing function-only handling to any named global
value and makes the addrspacecast operand representable.
Mirror both attributes in ModuleTranslation::convertMetadataAttr so the
[6 lines not shown]
[InstCombine] Fold shl of constant by cttz into multiply of lowest set bit (#214517)
Currently, `C << cttz(X, true)` generates a DeBruijn lookup table on
RV64I (13 instructions).
And this patch adds a fold in InstCombine:
`C << cttz(X, true) --> (-X & X) * C`
This reduces the instruction count from 13 to 3 on RV64I.
The fold requires that cttz has a single use (to avoid increasing
instruction count)
Alive2 proof: https://alive2.llvm.org/ce/z/TmWxrT
clang/AMDGPU: Stop passing redundant -target-cpu to cc1
Now that the exact target is encoded in the triple's subarch field,
-target-cpu is redundant. This avoids polluting the resultant IR with
unwanted "target-cpu" attributes. The net result is the desired codegen
when compiling libraries for a major subarch and linking it into a
program compiled for a specific arch. e.g., compiling for "gfx9-generic"
would pollute the IR with "target-cpu"="gfx9-generic", so codegen
would ultimately be performed for the generic target even after
linking into the concrete gfx9 cpu. The specialization will now be
achieved by merging the triples without the linker or optimization
passes needing to fixup function attributes.
AMDGPU: Start using subarch in attributor instead of subtarget
Avoid querying the subtarget for functions when the relevant
properties are known from the triple. The various subtarget
group size functions should also be decoupled from the subtarget,
but those are trickier to untangle.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
clang: Start using new amdgpu subarch triples
Fixup invocations using --target=amdgcn + -mcpu to introduce
the subarch in the triple.
For offload toolchains, a single toolchain is constructed for the
top level amdgpu architecture, and the effective triple is used for
target specific tool invocations.
The specifics of the resource directory layout are tbd. This does
try to find resources in the subarch named directory. The paths
are searched at toolchain creation time, so that does not work
when there are multiple subarches.
Fixes #154925
[Clang] Support friend declarations with a dependent nested-name-specifier (#208345)
Fixes https://github.com/llvm/llvm-project/issues/104057
---
This patch adds support for friend declarations with a dependent NNS
AMDGPU: Start using subarch in attributor instead of subtarget
Avoid querying the subtarget for functions when the relevant
properties are known from the triple. The various subtarget
group size functions should also be decoupled from the subtarget,
but those are trickier to untangle.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
clang/AMDGPU: Stop passing redundant -target-cpu to cc1
Now that the exact target is encoded in the triple's subarch field,
-target-cpu is redundant. This avoids polluting the resultant IR with
unwanted "target-cpu" attributes. The net result is the desired codegen
when compiling libraries for a major subarch and linking it into a
program compiled for a specific arch. e.g., compiling for "gfx9-generic"
would pollute the IR with "target-cpu"="gfx9-generic", so codegen
would ultimately be performed for the generic target even after
linking into the concrete gfx9 cpu. The specialization will now be
achieved by merging the triples without the linker or optimization
passes needing to fixup function attributes.
clang: Start using new amdgpu subarch triples
Fixup invocations using --target=amdgcn + -mcpu to introduce
the subarch in the triple.
For offload toolchains, a single toolchain is constructed for the
top level amdgpu architecture, and the effective triple is used for
target specific tool invocations.
The specifics of the resource directory layout are tbd. This does
try to find resources in the subarch named directory. The paths
are searched at toolchain creation time, so that does not work
when there are multiple subarches.
Fixes #154925
AArch64: Use MIPatternMatch in PostLegalizerCombiner ext checks (#216512)
Replace the getVRegDef + opcode-check idiom with mi_match. Add an
m_GSExtInReg matcher for G_SEXT_INREG, which has an extra immediate
operand and so does not fit the plain unary-op matcher shape.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
X86: Fix optimizeCompareInstr crash on a compare from an undef register
getVRegDef returns null for undef sources, so the assert would fire.
Found by AI while working on other stuff.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
PeepholeOpt: Fix crash on copy from an undef register
Fix ValueTracker looking at the def chain of an undef subregister.
Found by AI while working on something else.
Co-authored-by: Claude (Opus 4.8) <noreply at anthropic.com>
[ASan] Register assert lit feature in `lit.cfg.py` (#216529)
Add this feature to test a functional change related to asserts, see
https://github.com/llvm/llvm-project/pull/213546#discussion_r3790705838.
And this will allow other people to add `// REQUIRES: asserts` in
testcases in the future too.
I follow the same pattern already used in `flang/test/lit.cfg.py` and
other configs.
[SSAF][UnsafeBufferAnalysis] Address follow up questions after the approval of #209354
- The analysis should not create entries for empty contributors, which otherwise is non-empty in the serialized format.
- use std IO instead of a tmp file for regex-ing FileCheck queries.
rdar://179151541 & rdar://179151882