[BOLT] Default heatmap block sizes to cache line, pages and hugepage
The defaults were 64, 4K, 256K. 4K is the page size only on x86-64 and on
AArch64 kernels built that way; AArch64 also runs 16K and 64K base pages, and
256K corresponds to nothing in particular on either.
Use 64, 4K, 16K, 64K, 2M: the cache line, the three base page sizes in use, and
the PMD hugepage above a 4K base page. Each granularity then maps onto a real
capacity, which is what makes the working set numbers comparable to one -- L1i
lines, iTLB and L2 TLB entries, frontend region-table entries.
Two more granularities cost two more passes over an already-built map, no extra
decoding.
[BOLT] Keep the spelling of each heatmap block size
The block-size parser turns "64K" into 65536 and discards the original text,
keeping it only for error messages. The working set log then has to either
reprint the raw value or reformat it back, and reformatting invents a spelling
the user did not choose: "1MiB" comes back as "1M".
Store the spelling next to the value and echo it. Heatmap file names keep using
the numeric value, matching the existing "dumping heatmap with bucket size N"
message and the -<size> suffix that tests already expect.
Test Plan:
updated heatmap-preagg.test
[BOLT] Report working set size from the heatmap (#215429)
The heatmap produces a CDF (code coverage @ given sample pct), but only
writes it out as a table. Add two things:
1. Recompute CDF with scaled bucket sizes,
2. Log CDF at given sample pct (`-heatmap-cdf-pct` default p99) + total.
This effectively reports the code working set size (p99 and total),
expressed in units that map to uarch sizes: cache line (64B), base page
(4/16/64K), region table (2M), huge page (2M for PMD w/4K base), etc.
Sizes of interest can be specified using `-block-size=size1,size2,...`
Test Plan:
updated heatmap.test
wpa: Update to 2.12
Fixes and new features include:
hostapd:
* support RSN overriding (e.g., WPA3-Personal Compatibility Mode)
* EHT/IEEE 802.11be/Wi-Fi 7
- more complete support
- fix message validation issues that could enable DoS attacks
- fix group key rekeying
* enable SAE group 20 by default if SAE-EXT-KEY is enabled
* reject unexpected SAE password identifier to avoid DoS attack against
a specific STA
* mandate use of SAE H2E when using password identifiers
* assign VLAN when using SAE with PMKSA caching
* support SPP A-MSDU negotiation
* support IEEE 802.11bi functionality
- changing SAE password identifiers
- EPPKE
[50 lines not shown]
[lit] Stop bare env from short-circuiting the pipeline (#214512)
env with no trailing subcommand returns early from _executeShCmd,
skipping the rest of the pipeline. A RUN line like env | FileCheck never
runs FileCheck.
Runs it as an in-process pipeline stage instead, reusing the existing
InProcessPipe implementation.
Fixes #115578.
Also fixes six Inputs/shtest-env-positive fixtures that could never have
passed once FileCheck actually ran. Five expected KEY = VALUE from
FileCheck while env has always printed KEY=VALUE. Two RUN lines used a
check-prefix that matched no CHECK line in the file. One needed
--allow-empty since env -i produces no output at all.
release/Makefile.gce: migrate gsutil usages to gcloud CLI
Google Cloud recommends migrating from gsutil to gcloud storage CLI.
Update gce-do-upload target to use `gcloud storage buckets create` and
`gcloud storage cp` instead of `gsutil mb` and `gsutil cp` commands.
PR: conf/297016
(cherry picked from commit 4174cc2f69d36105a735b19fadc9c18497b02b1a)
release/Makefile.gce: migrate gsutil usages to gcloud CLI
Google Cloud recommends migrating from gsutil to gcloud storage CLI.
Update gce-do-upload target to use `gcloud storage buckets create` and
`gcloud storage cp` instead of `gsutil mb` and `gsutil cp` commands.
PR: conf/297016
(cherry picked from commit 4174cc2f69d36105a735b19fadc9c18497b02b1a)
[ORC] Standardize dylib manager on NativeDylibManager names (#215441)
The in-tree SimpleExecutorDylibManager now publishes a single controller
interface -- the ORC runtime's NativeDylibManager symbol names -- which
EPCGenericDylibManager targets, dropping the parallel LLVM-style
SimpleExecutorDylibManager_* names and the Create path that used them.
Follow-up to the EPCGenericDylibManager proxy refactor, which already
routed both name sets through the same code.
Details:
* Remove EPCGenericDylibManager::CreateWithDefaultBootstrapSymbols.
SimpleRemoteEPC and lli now construct via Create(ExecutionSession&),
which resolves the NativeDylibManager names in the session's bootstrap
JITDylib.
* Remove the SimpleExecutorDylibManager{Instance,Open,Resolve}
bootstrap-name constants from OrcRTBridge. SimpleExecutorDylibManager
previously published both name sets and now publishes only the
[4 lines not shown]
[DWARFCFIChecker] Changing register initial status as undefined (#209032)
As per DWARF spec 6.4.1. Users can overwrite this default assumption inside the prologue.
[mlir][SPIR-V] Fix swapped select operands in index.floordivs lowering (#214770)
Fix ConvertIndexFloorDivSPattern to select `negRes` when `cmp` is true
Previously operands were reversed, producing the wrong sign for the
floordiv result whenever the negative result branch should've been taken
[CIR] Drop the call-conv lowering flag from the new CodeGen tests
The two CodeGen tests this branch adds ask for the pass with
`-clangir-enable-call-conv-lowering`, which no longer exists. The pass
now runs by default, so the RUN lines do not need a flag at all and get
the same lowering without one.
Assisted-by: Cursor / claude-opus-5
usb: xhci: allow up to 1s for SET_ADDRESS
Some devices take a little longer, and the spec doesn't really seem to
mandate a maximum. The common path in usbd_req_set_address() has
already been bumped to 1s and I have a headset (Logitech H390) that does
need a little bit longer, so let's match it in xhci.
Reviewed by: aokblast
Differential Revision: https://reviews.freebsd.org/D58717
[CIR] Enable x86_64 calling-convention lowering by default (#215026)
x86_64 calling-convention lowering has been opt-in behind
`-clangir-enable-call-conv-lowering` since it landed, so nothing reaches
the pass unless a test asks for it. The ClangIR default path therefore
emits high-level signatures that do not match SysV: a 32-byte struct
return stays first-class instead of going out through `sret`, and a
one-eightbyte struct argument is passed as a record instead of being
coerced to `i64`.
This change turns the pass on by default for x86_64 and renames the flag
to a `BoolFOption` pair, `-fclangir-call-conv-lowering` and
`-fno-clangir-call-conv-lowering`. The last flag on the command line
wins, so a build can disable the pass globally and re-enable it for one
translation unit.
Flipping the default breaks 147 of 949 CIR tests. 37 tests are CHECK
regenerations where lowering moved toward classic CodeGen. 110 tests are
quarantined by adding the disable flag to the RUN lines that turn the
[12 lines not shown]
[SimplifyLibCalls] Shrink llvm.sincos.f64 to llvm.sincos.f32 (#211218)
The double -> float shrink in LibCallSimplifier::optimizeCall only knows
about Intrinsic::sin and Intrinsic::cos, so once InstCombine combines a
sin/cos pair into llvm.sincos the fpext/fptrunc pair is left in place
and the work is done at double precision.
Handle Intrinsic::sincos too, behind the same UnsafeFPShrink gate. It
needs its own helper rather than optimizeDoubleFP because sincos returns
a struct: the results are read back through extractvalue instead of
being used directly, and the narrowed call has to be rebuilt as a
struct. InstCombine folds the resulting extractvalue/insertvalue pairs
away, so the sin/cos pair in the reported case now ends up as a single
llvm.sincos.f32 call.
NOTE: reported as a 2-2.5% regression on SPEC17 526.blender with -flto
-ffast-math on neoverse-v2 in #194616.
Assisted-by: Opus 4.8
[libc++][CI] Add a cron job to trigger benchmarking jobs (#212859)
This patch introduces a GitHub workflow that runs on a schedule and
determines which benchmarking jobs to trigger to fill the LNT instances
with performance data.
As a drive-by, it also consolidates the information describing libc++
LNT machines into a single JSON file.
The added cron workflow will run every hour, but since the benchmarks
typically take more than an hour to run, it is expected that some of
those triggers will not actually trigger new jobs.