70 points surprisetalk 1 day ago 15 comments
vlovich123 2 hours ago | parent
> Anything faster than, say, 10ms risks being skewed by fixed costs (e.g, interpreter startup).
Sounds like the author’s experience is strictly in Python. For example with Java you have to make sure the JIT has sufficient optimized your program.
Additionally there’s plenty of situations where it can take a really long time to generate a representative dataset worth benchmarking and it can take time to evaluate the performance (eg databases). Short and quick microbenchmarks can be useful as building points, but at some point you need to evaluate steady state performance of the full thing. Other domains this comes up with is game rendering performance where a 300ms sample tells you nothing about whether you have frame drops after minute 25 or have a memory leak.
thadt 1 hour ago | parent
However in general - I agree with OP. Most of the time I care about milliseconds.
> Sounds like the author’s experience is strictly in Python. Er [1], no [2].
jonhohle 1 hour ago | parent
There also needs to be care taken in how these measurements are aggregated. Averages will almost always tell you nothing. High percentiles (95%, 99%, 99.9%) under load may show you something completely different than the average or even median case.
winwang 2 hours ago | parent
jsd1982 2 hours ago | parent
vardump 2 hours ago | parent
I tried to fix it by switching hyperthreading off, playing with the scaling governor, boost, setting a CPU frequency to no avail. The jitter was too much and the results were not reproducible, so I just gave up.
Of course your mileage may vary; this was on an AMD Zen 3 CPU.
hliyan 1 hour ago | parent
a) language was not garbage collected (C++)
b) we avoided heap lock contentions in critical paths by pre-allocating object pools at startup
c) I/O operations were offloaded to separate threads, connected by mutex locked linked lists
d) processing thread was bound to its own CPU core
That's about as deterministic as we could get.
vardump 14 minutes ago | parent
I think in 2008 CPUs were not so crazy about power and heat management.
veritron 2 hours ago | parent
hyperpape 1 hour ago | parent
The harder you push, and the more you need to start finding smaller improvements, the more this advice becomes a rule of thumb you can't rely on.
bhouston 1 hour ago | parent
[1] https://developer.mozilla.org/en-US/docs/Web/API/Performance...
bob1029 1 hour ago | parent
I also prefer microseconds when working with SQLite and AspNetCore.
spacedcowboy 1 hour ago | parent
So if I'm comparing against (say) C++, Swift, ObjC, on a given (arm64 or x86_64) architecture, I want to see that sort of timescale, a bit less is fine, a bit more is fine.
Of course, when you (today [grin]) get auto-vectorisation of 2D matrix multiplies, and you have an SME/SME2 target on arm64 that most compilers don't pick up so you're 150x faster than clang/g++, you might have to run it a bit longer so you can get reasonable comparison numbers :)
cchianel 1 hour ago | parent
- Does several warm-up runs so the JIT-optimized code is benchmarked instead of the interpreted code/compilation.
- Create multiple forks of the JVM to eliminate JVM run variance.
- Provide utilities like `Blackhole` to prevent dead code elimination and `State` to do setup and prevent constant folding.
spankalee 40 minutes ago | parent
It's very hard to say much about an absolute number. You need to compare against some alternative or control, and because of CPU load, throttling, GC, and a thousand other variables, you really should be comparing against that control _in the same run_, and importantly round robin across multiple runs to spread out the noise fairly across each implementation.
Then once you have a bunch of measurements you have a distribution and shouldn't just take a mean to compare, but should calculate something like the 95% confidence interval. If you see that those confidence intervals overlap, the you might not really know which is faster. If they don't overlap, then you probably do know which is faster.
If you have a good benchmark, then running it more times can narrow the confidence intervals and let you tease out very small improvements at the cost of longer runs. If confidence intervals don't narrow, then you hit the limits of signal-to-noise.
This is the only way I've been able to get reliable, actionable benchmarks outside of a very, very controlled hardware lab. It's what Google's Tachometer benchmark runner does, and I wish more runners did this: https://github.com/google/tachometer