[BOLT] Skip data-hole filling unless data reordering is enabled
BinaryContext::postProcessSymbolTable unconditionally called
fixBinaryDataHoles(), which walks every allocatable section and, for
each gap in its address space, either grows a zero-sized data symbol
or creates a synthetic "HOLEat" BinaryData (plus an MCSymbol and
GlobalSymbols/BinaryDataMap entries). This machinery was introduced
(0e4d86bf, 2017) for one purpose: to give static data reordering
(-reorder-data) a movable object covering every byte of a section. It
has no other consumer.
On a large binary, these synthetic objects are live from
buildFunctionsCFG through the end of the run and, at the RSS peak
(during debug info rewriting), fixBinaryDataHoles accounted for 1669
MB (2.5%) of peak RSS -- memory spent entirely for a feature that is
off by default.
Gate fixBinaryDataHoles() (and the zero-sized-symbol validation loop
that presumes it ran) on a non-empty opts::ReorderData, keeping
generateSymbolHashes() unconditional.
[BOLT] Key GlobalSymbols on MCContext-owned names to reduce memory (#214891)
BinaryContext::registerNameAtAddress registers every symbol name twice.
It first calls MCContext::getOrCreateSymbol(Name), which interns the
name in MCContext's symbol table (the MCSymbol owns the string via its
table entry). It then also stored the name in the GlobalSymbols map,
which was a StringMap<BinaryData *>. StringMap owns its keys, so each
global name was duplicated: one copy in MCContext and a second copy in
GlobalSymbols. Both grow with the number of symbols and, for large
binaries with long mangled names, this duplication is a meaningful
source of memory use during file object discovery.
This change makes MCContext the single owner of these name strings and
have GlobalSymbols merely reference them. GlobalSymbols becomes a
DenseMap<StringRef, BinaryData *> keyed on the MCContext-owned name
(MCSymbol::getName() of the symbol just created/looked up). No string is
copied into the map: each entry is a fixed-size (StringRef, pointer)
pair regardless of name length. Lookups (getBinaryDataByName, count) are
unchanged because DenseMap<StringRef> hashes and compares by content,
[7 lines not shown]
[AMDGPU] Correct DS FIFO buffer size semantics
There was some ambiguity in how buffersize 0 and 1 are handled. The
correct semantics are:
- `BufferSize == 0`: unlimited, no FIFO stall
- `BufferSize == 1`: unbuffered, only one instruction in flight
- `BufferSize > 1`: buffered FIFO
[BOLT] Page out .dwo files (#214903)
Split-DWARF inputs at big binaries scale ship 100+ GiB of .dwo files.
BOLT opened a fair number of them during readDebugInfo, putting a lot of
pressure on the OS memory management: mmap'd reads always populate the
page cache; with every .dwo mapped at once those pages accumulated,
refaulted, and registered as memory pressure that got the process
oomd-killed.
Now, .dwo page-cache pages are reclaimed as soon as BOLT is done with
each file: madvise(MADV_PAGEOUT) on the live mapping, then
posix_fadvise(POSIX_FADV_DONTNEED) once it is unmapped. Controlled by
-drop-dwo-page-cache, OFF by default, as it is unlikely upstream will be
processing gigantic sets of dwo files.
[BOLT] Create and release .dwo DWARF contexts incrementally (#214900)
BOLT opened a DWARFContext for every .dwo during
readDebugInfo and kept them all alive until teardown. On large
split-DWARF targets that is tens of GiB held resident through emission,
the point of peak RSS.
Make the DWOCUs map a lazily-populated cache instead:
* Use the newly added DWARFUnit::clearDWO()/hasDWO() to directly manage
DWARFUnit's DIE caching mechanism.
* BinaryContext::getDWOCU() opens a context on demand (keyed off a
stable DWOId -> skeleton CU map).
* Release contexts as soon as they are done with: all of them at the end
of readDebugInfo, and per-bucket at the DWARF rewrite merge point.
* Remove DWOCUs map, which became redundant and whose purpose can now be
served by the new id-to-skeleton map, and then fetching the split CU
from the skeleton via getNonSkeletonUnitDIE().
[3 lines not shown]
[CIR][NFC] Fix loop op examples in CIROps.td (#216457)
### summary
The cir.while example had cond and body swapped, and several loop
examples used outdated cir.condition / cir.for syntax. Also add short
examples of the optional per-iteration cleanup region.
Generated by Grok 4.6, but manually reviewed.
[BOLT] Page out .dwo files
Split-DWARF inputs at big binaries scale ship 100+ GiB of .dwo
files. BOLT opened a fair number of them during readDebugInfo, putting
a lot of pressure on the OS memory management: mmap'd reads always
populate the page cache; with every .dwo mapped at once those pages
accumulated, refaulted, and registered as memory pressure that got the
process oomd-killed.
Now, .dwo page-cache pages are reclaimed as soon as BOLT is done with
each file: madvise(MADV_PAGEOUT) on the live mapping, then
posix_fadvise(POSIX_FADV_DONTNEED) once it is unmapped. Controlled by
-drop-dwo-page-cache, OFF by default, as it is unlikely upstream
will be processing gigantic sets of dwo files.
[BOLT] Create and release .dwo DWARF contexts incrementally
BOLT opened a DWARFContext for every .dwo during
readDebugInfo and kept them all alive until teardown. On large
split-DWARF targets that is tens of GiB held resident through
emission, the point of peak RSS.
Make the DWOCUs map a lazily-populated cache instead:
* Use the newly added DWARFUnit::clearDWO()/hasDWO() to directly
manage DWARFUnit's DIE caching mechanism.
* BinaryContext::getDWOCU() opens a context on demand (keyed off a
stable DWOId -> skeleton CU map).
* Release contexts as soon as they are done with: all of them at the
end of readDebugInfo, and per-bucket at the DWARF rewrite merge
point.
* Remove DWOCUs map, which became redundant and whose purpose can
now be served by the new id-to-skeleton map, and then fetching
the split CU from the skeleton via getNonSkeletonUnitDIE().
[5 lines not shown]
[DebugInfo] Add DWARFUnit::clearDWO() (#214899)
Add DWARFUnit::clearDWO() so a skeleton unit can drop the DWO context it
owns without being destroyed itself. Also add DWARFUnit::getDWO() to
answer, without creating a new one, if that skeleton CU is currently
caching a DWO context.
For example, BOLT opened a DWARFContext for every .dwo during
readDebugInfo and kept them all alive until teardown. On large
split-DWARF targets that is tens of GiB held
resident. clearDWO()/getDWO() expose to users DWARFUnit's caching
capacity, allowing them to spontaneously drop the cache/look it
up/re-load it for memory management. To demonstrate this, in
llvm-dwarfdump we now make use of the same technique to avoid ballooning
peak RSS when dumping binaries with dwos. Whenever --debug-info --dwo is
used, during the loop dumping non-skeleton DIEs, we clearDWO as soon as
we're done with that unit. Testing on a large binary, this was shown to
reduce peakRSS from 50GB to 590MB.
Handle DS FIFO accounting edge cases
Saturate hardware-unit pressure decrements and treat buffer sizes zero
and one as disabling buffering to avoid underflow and inconsistent stall
costs.
internal/resources: Expand the commit bit and resources policy
Elaborate on how the core team delegates management of
repository-specific resources, such as commit bits, to the responsible
teams.
Add guidelines for suspending or revoking a commit bit: except where
circumstances justify immediate action, teams are expected to contact
the committer and provide an opportunity to address the problem before
acting, and to inform both the committer and the core team when a commit
bit is suspended or revoked.
Reviewed by: core (adrian, imp)
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D58830
Update perl to 5.42.3
We already had the security changes, this is mostly version bumps
and documentation adjustments for them. There does pull in additional
changes in Archive-Tar, Compress-Raw-Bzip2, and IO-Compress.
While here, bump the libperl version for the previous CVE that
changed a public header.
OK and suggestions bluhm@
[AMDGPU] Use DS latency for FIFO scheduling
Use instruction latency for DS hardware-unit cycle accounting so the FIFO
model can identify a full buffer. Add focused MIR coverage for the resulting
stall cost and scheduling decision, and regenerate the integration checks.
Change-Id: I2f4df2e97d145af4935872dbd43108e1b55077ab
linux: add dma-buf and sync_file ioctl handlers
drm-kmod already implements the dma-buf and sync_file ioctls, but
linux_ioctl.c had no handler group for the 'b' and '>' magic bytes, so
the requests never reached it and returned EINVAL from
linux_ioctl_fallback(). Route the commands drm-kmod services to
sys_ioctl(), translating the direction bits with SETDIR(); everything
else still falls through to the fallback and keeps getting named in
dmesg.
Approved-by: adrian
Accepted-by: dumbbell
Signed-off-by: Nick Price <nprice at FreeBSD.org>
(cherry picked from commit d6a7e89504af337413af39fd121026f512c0a35d)
[orc-rt] Use StringOutputStream in executor error messages (#217167)
Replace std::ostringstream and std::to_string in the executor's
error-message construction with orc_rt::StringOutputStream, dropping
<sstream> (and its iostreams + locale machinery) from the executor,
which runs on-target including on freestanding platforms.
Converted SimpleNativeMemoryMap.cpp, ExecutorProcessInfo.cpp, and
Unix/NativeDylibAPIs.inc; NativeDylibManager.cpp no longer needs to
include <sstream> for the latter. Messages that only feed an error are
built from a StringOutputStream temporary; one all-literal message drops
the stream entirely.
Hex values now carry a "0x" prefix (StringOutputStream::hex always emits
one), where the old std::hex produced bare digits. No test depends on
the exact message text.