[mlir][ROCDL] Add TargetInfo to replace Chipset, allow features queries (#223562)
Add a new ROCDL::TargetInfo struct that parses AMDGPU triples and
target names using the same logic that Clang and LLVM
use (TargetParser) and maintains the set of features available on a
given GPU.
This is an improvement over the old `amdgpu::Chipset` struct since
that was just a version number and often became stale compared to the
knowledge exposed by LLVM, such as gfx1170 having OCP FP8 support even
though other gfx11 chips don't have it.
This struct also allows for moving to new-style
triples (amdgpu9.42-amd-amdhsa vs amdgcn-amd-amdhsa--gfx942, for
example), which is an ongoing migration in other parts of the compiler
that this PR lets us follow.
It also enables compiling for generic targets, like `gfx11-generic`,
which can be run on all chips in a generation.
[14 lines not shown]
[AMDGPU] Canonicalize num_records to its actual width in InstCombine (#217068)
This PR adds code to InstCombineIntrinsic to change the width of the
num_recods field (by extension or truncation) to the correct width for
the target triple (if a concrete enough target triple has been set) so
that LLVM IR-level optimizations can see the lack of demand for the
high bits, for example.
Assisted by Claude, which also found those buffer lowering edge cases
Docs: remove gdb startup text from WritingAnLLVMPass. (#228236)
It is simply noise for readers of the doc.
Also, it is (incorrectly) flagged by license-compliance scanners as
indicating GPL-licensed content. Simplest to just drop it.
mail/cone: Update to 2.5
Drop pkg-deinstall:
@sample already removes etc/cone on deinstall when unmodified. The
script ran before the files were removed, so it always reported the
config as modified and left in place, even when pkg then deleted it.
ChangeLog:
https://sourceforge.net/p/courier/courier.git/ci/master/tree/cone/ChangeLog
[GlobalISel] Reject scalar sources in vector truncate combine (#228121)
G_UNMERGE_VALUES may have a scalar source. Calling getNumElements() on
its LLT asserts, so reject scalars before checking the vector element count.
Add a scalar regression to the existing AArch64
prelegalizer-combiner-use-vector-truncate.mir test. Its existing vector
combine cases remain unchanged.
[LV] Add tests for narrowing interleave groups with factor > VF (NFC). (#228217)
Add initial test cases with opportunties to narrow interleave groups
where number of members > VF.
[WinEH] Emit unreachable on malformed catchpad (#222181)
Fixes #219223 by protecting against malformed catchpad arguments and
marking all catchpads under the parent catchswitch as unreachable.
The catchpad arguments contain critical information for the exception
handling runtime, which I understand gets baked into the binary's
exception handling tables. We can't really put "nothing", and any value
we do put would corrupt the whole EH table. That is why we need to
disable all the associated catchpads, not just the malformed one.
[LV] Add tests for narrowing interleave groups sharing wide ops (NFC). (#228209)
Add tests where wide recipes feeding store interleave groups are used by
multiple groups or with different members. Currently
@multiple_store_groups_sharing_member_0 and
@wide_op_used_with_different_members are miscompiled with VF 2.
[flang][OpenMP] Improve check for LINEAR and ORDERED with argument
The LINEAR clause is not allowed on a construct if an ORDERED clause
with an argument is present. The previous check rejected it even when
the ORDERED clause had no arguments.
[flang][cuda] Reject CONSTANT in main program (#228193)
From the CUDA Fortran programming guide:
> In the section on variable qualifiers: "The constant variable must be
declared within a module's global data specification scope. … All host
accesses of constant memory must be through use or host association."
> In section 3.2.5, Constant data: "Constant variables and arrays can
appear in modules … Constant variables appearing in modules may be
accessed via the use statement in both host and device subprograms." The
list of allowed uses in host code includes "a named entity within a USE
statement" and "a dummy argument in a host subprogram", but not a local
declaration.
[lldb-dap][test] Do not overwrite failed/error logs. (#227857)
On failure, log files name changes to be unique to a test method. Append
testcase and DAP logs to log_files so it doesn't get overwritten by the
next test method.
[AMDGPU] Validate scale_sel in v_cvt_scale_* (#227475)
These instructions can be block16 or block32 depending on the target
and scale_sel bits. Block16 is not supported in strict mode.
Re-enable the rest of the instructions in the strict mode but validate
the scale selector.
[Polly][Test] Fix missing -plugin-arg (#228221)
After #226773, plugin options must be passed via -plugin-arg. Fix the
regression test from #227311 to the changed command line.
[lldb][Windows] Report a process exit once the process is gone (#228115)
When a debugged process exits, lldb and lldb-server report the exit
while Windows is still tearing the process down. Its executable and DLLs
are still loaded at that point, so a test that deletes or replaces one
of them right after the exit fails with `"access denied"`. This is what
makes `TestReplaceDLL` fail on Windows CI.
This patch reports the exit only after the debugger has released the
exit event and the process has fully terminated, which matches how POSIX
reports an exit.
On a Windows 11 host, `TestReplaceDLL` fails to delete `foo.dll` in 246
out of 640 runs before the change and in 0 of 1,600 after.
rdar://188918066