[CIR] Add structured control flow for coroutine suspend points (#213191)
This PR adds 2 new ops for improving structured control flow in
coroutines, according to the redesign in this discussion:
https://github.com/llvm/llvm-project/pull/203802#discussion_r3464356343
One of the ops is `suspend_point`, which marks the point where a
coroutine should suspend and return control flow to the caller it's also
the point where the coroutine resumes from when it starts again. This op
must be inside the `cir.await` suspend region, and it works in pair with
the `CoroRetPoint` op, which marks the only exit point of the coroutine
so no matter how many awaits/suspend points exist in the body, they all
converge on the same exit path.
This will help the FlattenCFG pass know where to branch when flattening
the `cir.await` op. Previously we didn't have an exact point to branch
to when suspending a coroutine this gives us that.
---------
Co-authored-by: Erich Keane <ekeane at nvidia.com>
[CIR][OpenMP] Add support for combined target parallel directives (#207019)
This patch adds support for handling combined target and parallel
directives. It also refactors the code so that the emission functions
can be easily combined to handle both the single and combine cases for
emitting directives.
Assisted-by: Cursor / claude-opus-4.8-medium
[OpenMP][libomp] Cleanup OpenMP specification version information (#225427)
Make OpenMP specification version and date aligned with what is defined
in the CMake configuration.
fifo: Fix EOF reporting in various cases.
According to POSIX:
> When attempting to read from an empty pipe or FIFO:
>
> If no process has the pipe open for writing, read() shall
> return 0 to indicate end-of-file.
So:
1. When there are no writers, after open(O_RDONLY|O_NONBLOCK),
blocking reads should immediately report EOF instead of blocking,
and should continue to immediately report EOF on repeated reads.
Previously, they would simply block, because the wrong socket was
initialized with SS_CANTRCVMORE -- though curiously, nonblocking
reads would consistently report EOF.
[42 lines not shown]
databases/postgresql-postgis2: Drop explicit ACCEPTED
postgis builds with all versions in the default value of
PGSQL_VERSIONS_ACCEPTED. It's generally not difficult about pgsql
versions, and likely to continue to work with all pkgsrc versions.
Drop the explicit redundant setting as unnecessary, as well as
confounding for those adjusting mk/pgsql.mk to only build one version.
[AMDGPU] Omit hardwired-on SRAMECC ELF mode
Mirror #227740 for SRAMECC. Without on/off modes, SRAMECC is implied
by EF_AMDGPU_MACH; encoding a mode makes consumers infer an
unsupported :sramecc+ modifier.
Depends on #225540.
Change-Id: I7bf76efbb86c7030eae36dc06ca7da0dd64cce87
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[AMDGPU] Add SRAMECC on/off mode capability (#225540)
Distinguish SRAMECC hardware support from selectable on/off modes,
mirroring XNACK. Preserve selectable modes on existing targets and use
the new capability for target-ID validation, defaults and printing,
module flags, and disassembly.
sendmsg(2): Report ENOBUFS instead of success on failure to pass fds.
PR kern/60832: AF_LOCAL stream: sendmsg() with SCM_RIGHTS silently
drops data and descriptors but reports success
cmsg: Test edge case of fd-passing near buffer size.
This test allows sendmsg to fail with ENOBUFS without blocking, and
verifies that the fds are received if it succeeds.
When passing file descriptors, sendmsg can fail with ENOBUFS it did
not or would not block because _after_ the blocking criterion is
tested in the AF-generic uipc_socket.c logic (essentially, whether
the data length + control length would exceed the send buffer size),
the control buffer is expanded on some architectures by converting
each int file descriptor to a struct file pointer in the kernel, and
then the buffer size is checked again in AF_LOCAL-specific logic when
unp_send calls sbappendcontrol.
Frankly I think this is a bad design, and sendmsg should just not
fail for this reason if it has passed the blocking criterion: either
(a) the AF_LOCAL-specific logic should count the user's control
buffer size with ints rather than the the kernel's control buffer
[16 lines not shown]
[SanitizerCoverage] Don't let noipa affect comdat selection (#228182)
On non-ELF targets, SanitizerCoverage only puts a function's coverage
arrays in a comdat if the function isn't interposable.
`isInterposable()`
also returns true for `noipa` function definitions by default, but
`noipa` doesn't affect linkage. Query
`isInterposable(/*CheckNoIPA=*/false)` so `noipa` functions get the same
comdat treatment as other functions with the same linkage.
This is one of a few refinements of noipa identified while working on
making optnone imply noipa - without these refinements,
optnone-implies-noipa, might substantially change clang -O0 codegen. I'm
open to discussing whether these refinements are the right direction,
though.
Assisted-By: Claude
[DebugInfo] Support global addresses in variable locations (#218722)
This patch adds support for global address locations to SelectionDAG,
GlobalISel and AsmPrinter. The actual implementation is very similar to
how constants are being handled so most code changes are very
mechanical.
The motivation for adding global address support is the Swift compiler
and it is needed even in unoptimized code. For example, the standard
libary implements the empty array with a global "empty array storage"
singleton object, so it can show up as the location of array values.
Another very common use-case are references to type metadata, which is
also accessed via global symbols.
It is even possible to motivate this with a C exmple:
extern int g; int *p = &g;
also produces the dbg_value with a global value.
[17 lines not shown]
[compiler-rt] Remove dlsym interceptor and support `-shared-libsan` for CSan
Summary:
Follow the UBSan offload runtime. Offload now resolves HSA through the
global scope, so the `dlsym` interceptor is no longer needed. The real
HSA entry points are still taken from the loaded HSA library rather than
`RTLD_NEXT`, since every DSO with a static runtime exports the same
wrappers and they would otherwise chain back into each other.
Build `libclang_rt.csan.so` with the offload objects folded in. The
exported HSA wrappers report failure when HSA is absent and warn when HSA
was loaded ahead of the runtime. The preinit hook moves to a separate
`csan_offload-preinit` archive for executables.
[Clang] Support `-shared-libsan` for offload CSan
Summary:
Follow the UBSan handling. The shared runtime embeds the HSA
interceptors, so `csan_offload` is only linked with static runtimes and
executables pull in `csan_offload-preinit` to initialize early.
[compiler-rt] Add 'csan' library for the concurrency sanitizer
Summary:
Adds the runtime for the concurrency sanitizer, both CPU and GPU.
Fundamentally, this works using the following pseudocode:
```c
static u64 watchpoints[N]; // Hash-indexed, zero is empty.
// Emitted before the access, so we never trip on our own write.
void check_access(volatile void *addr, u32 size, u32 type) {
// Every access probes. A read conflicts only with a watched write, a
// write conflicts with either.
if (u64 *wp = find_watchpoint(addr, size, type))
consume(wp, this_pc()); // Hand our location to the owner.
if (!should_sample()) // Wave-uniform, 1-in-N chance.
return;
[17 lines not shown]