Repository navigation
perf: measure against the alternatives, and fix what that found - #90
Merged
Merged
Conversation
The path list was maintained by hand, and a hand-maintained list stops covering a new directory without anyone noticing -- benchmark/compare was never checked. Worse, an entry whose glob matches nothing does not expand to nothing: bash passes the literal pattern through and clang-format fails on it, so adding a path for a directory that does not exist yet breaks the job. git ls-files avoids both. The .in templates stay out on purpose: they hold CMake @substitutions@ and are not valid C++ until configured. Verified: the discovered set is the same 13 files with zero violations, and a deliberately misformatted file in a directory the old list did not cover fails the check. Note that git ls-files reports tracked files only, so a brand-new file is checked once it is staged, not before -- which is what CI sees anyway, since everything is tracked after a checkout.
The 41 micro-benchmarks in benchmark/src measure single operations against nothing. The pitch has always been implicitly about speed -- static sizes, always_inline, tag types -- with no comparative number anywhere behind it. This adds a Compare target that measures what a caller wants: the derivatives of one non-trivial function, against the libraries someone choosing between them would weigh. The function is the strain energy of a cable with unit-spaced nodes, so each term is the squared elongation of a segment; it is nonlinear, its Hessian is not constant, and it is defined for any n. Two separate comparisons, because the orders are separate questions. Second order (HyperJet static and dynamic, Eigen AutoDiffScalar nested, autodiff dual2nd, central differences): HyperJet is fastest everywhere, but by 1.3x to 1.7x from n=6 up, and at n=3 nested AutoDiffScalar is level. The big factors are against finite differences and against HyperJet's own dynamic variant, which costs 36x the static one at n=3. First order (the same plus ceres::Jet): HyperJet does not win. Up to n=12 the top three are level; at n=24 Eigen is 2.4x faster and Ceres 1.9x. Decomposing that case rules out the setup -- building the variables costs 97 ns against Eigen's 73 -- and points at the evaluation: 513 ns against 182. Roughly 0.6 double-operations per cycle against Eigen's 4.8, which looks like the first-order derivative loops not vectorizing where Eigen's expressions do. Recorded as a defect to fix, not written around. ceres::Jet is first order only, which is what a solver needs -- Jacobians. Nesting Jet inside Jet to force it into the second-order table would benchmark a configuration no Ceres user writes, so it is not there. jet.h needs only its own include tree and Eigen, so the dependency is one DOWNLOAD_ONLY package and no solver build. Three things the benchmark had to get right to be worth anything. It verifies every contender against HyperJet before timing -- a fast contender computing the wrong thing would read as a win. That earned its place immediately by catching Eigen's Hessian coming out as exactly zero. The cause was mine: both AutoDiffScalar and autodiff's dual are expression-template types, so deducing intermediates with auto kept expressions pointing at operands that were already gone. Every intermediate is now spelled out as the scalar type. Filling a fresh Result inside the timing loop put three vector allocations in every iteration and the n=3 second-order figure at 77 ns instead of 18 -- three quarters of it was std::vector, not differentiation. The caller allocates once now. And the loop guards against hoisting, the input being fixed and the computation pure. Measured both ways: hoisting was not happening here, the guard stays because relying on that is not a plan. The justfile now discovers its file list the same way CI does, so the two cannot drift apart again.
oberbichler
force-pushed
the
bench/compare-against-alternatives
branch
from
July 28, 2026 14:57
b618e22 to
b55c5c8
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Performance improvements and benchmarks