Hacker News new | ask | show | jobs
by sade_95 26 days ago
What My Project Does

GP_ELITE is a symbolic regression engine in pure Python: given (X, y) data, it searches for a readable mathematical formula linking them, instead of a black-box model.

To show what that means concretely: I gave it nothing but the 8 planets' distance from the Sun and orbital period — 8 data points — and asked for a formula. It returned:

T = a · sqrt(a) (i.e. a^1.5), R² = 1.000000

That's Kepler's Third Law (T² ∝ a³), which took Kepler ~10 years to find in 1618. GP_ELITE found it in ~3 seconds. Reproducible: examples/kepler_demo.py.

v0.2.0 (this week) added the parts that make it reliable: Levenberg-Marquardt constant fitting (constants come back at machine precision — Coulomb's q1·q2/(4πεr²) is recovered exactly), multi-restart with a merged candidate archive, a Pareto front output (the full complexity ↔ accuracy staircase, not just one champion), and a guarded forecasting mode for extrapolating trends beyond your data without the usual GP blow-ups.

Pure Python/NumPy — pip install gp-elite, no compiler, no Julia.

Target Audience

Anyone with small experimental datasets (≤10 variables, 100–5000 points) who wants to understand a relationship, not just predict it: lab engineers, scientists, students. One concrete use case that drove development: battery degradation (SOH) forecasting — the guarded mode gives you an honest bracket of scenarios (a Pareto front from a conservative straight line to richer bounded laws) instead of one overconfident curve. Production-usable for that niche (built-in hold-out validation, regression-tested); not aimed at large-scale ML.

Comparison

vs gplearn (the established pure-Python option): I ran both on the same frozen benchmark — 15 Feynman physics equations, identical data and splits, generous budget for gplearn. Exact symbolic recovery (machine precision): GP_ELITE 10/15 (67%) vs gplearn 6/15 (40%). gplearn recovers the constant-free formulas and stalls as soon as a ½ or a 4π appears (no real constant optimization); LM fitting is what closes that gap. Every number is reproducible: PYTHONHASHSEED=0 python benchmarks/feynman_bench.py 0 15 and benchmarks/duel.py in the repo.

vs PySR / Operon (the state of the art): they are stronger on speed and scale, and I'm not claiming otherwise — but they require a Julia or C++ toolchain. GP_ELITE's whole point is zero barrier: pip install and go.

vs neural nets / gradient boosting: those win on raw accuracy for large data, but give you a black box — GP_ELITE gives you the actual equation.

Honest limits: weak on chaotic targets (tested on Collatz), degrades past ~6 variables with decoy features, and pure Python costs wall-time on big data.

Code (MIT): https://github.com/ariel95500-create/gp-elite

1 comments

This is not surprising at all and depends on the inductive bias hardcoded in the search.

There are infinite number of curves that agree on those 8 points and deviate from Kepler 's law everywhere else. On such 'trajectories' this algorithm would have performed badly.

You're right, and I'd go further: the inductive bias isn't incidental, it's the whole product. Short trees over {+, *, sqrt, exp, ...} plus a parsimony penalty is basically Occam's razor made executable. A bias-free learner can't generalize at all (no free lunch), so the honest question is whether this particular bias matches the domain. For physics it has an unreasonably good track record though that's the mystery of physics, not of my library. Two nuances though. The evidence here isn't "a curve fits 8 points" infinitely many do, as you say. It's that a 3-node formula fits them to machine precision (1−R² ≈ 1e-15). Under an MDL view that's not nothing: the probability that such a short description nails 8 independent points exactly, if the truth were some unrelated wiggly curve, is astronomically small. The shortness is the evidence. Second: your adversarial curve would fool Kepler too, and any finite-data method ever. The practical mitigation is the boring one held-out validation, and in the planetary case, extrapolation: the law found on 8 planets keeps working on moons, asteroids and exoplanets. On the tool side I try to keep the failure mode visible rather than hidden: it returns the full accuracy-vs-complexity Pareto front, and the docs say plainly that noise breaks symbolic recovery long before it breaks fit quality. So yes: it finds simple laws when simple laws exist. When they don't, it fails ideally loudly.
I would argue that Kepler's success influenced the choice of the inductive bias. In that case the claim that Kepler took years what this can do in seconds is not an unbiased position to take.

I do see a great value in conjecturing possible solutions that needs to be verified with domain specific knowledge.

Recall that Ptolemaic epicycles were a great fit, in fact a better fit than Copernicus's heliocentric model. This makes me wary of deep NNs in Physics.

You are right, and I concede the point. I chose the operators and the parsimony rule knowing already that laws like Kepler exist. So "Kepler took 10 years" is a bit of theater, the honest claim is only: if a short law exists in this basis, the search finds it fast. Kepler had to invent the idea that such law exists at all, and fight Tycho's data with no computer. Not the same job.

For the epicycles, for me this is exactly the argument for parsimony pressure. Epicycles fit better because you can always add one more circle, same as adding parameters in a NN. They lose on description length, and that is the only defense I have too. And it is not perfect: on noisy data my tool produces its own epicycles, formulas that fit very well and mean nothing. One battery model it gave me predicted the battery heals itself after cycle 264. Great fit, zero physics. So yes, a conjecture machine that a domain expert must verify. No more than that.

I agree and I vehemently share your concern about profusion of tweakable parameters in the model.

There is some misconception in the wild about epicycles models that need not be shared by you specifically. There weren't that many epicycles per orbiting body, but every orbiting body had a few that had to be 'trained' specifically for them.

My fear, and I suspect yours too is that good curve fits done one at a time with mathematical models that are universal approximators (*) rarely, if at all lead to causally explanatory models. In Physics it's the latter that we seek.

(*) Epicycloidal models are a form of Fourier analysis and are a class of universal approximators for periodic trajectories.

Thanks for the epicycles clarification, I did not know each body had its own fitted set. This makes the parallel with overfitting even stronger.

And I think we agree on the real problem: a good fit on one problem, made with a universal approximator, is not an explanation. It can be just a compressed description, maybe with a cause behind, maybe not. I cannot fix this, but I added this week a small thing to at least see it: a bootstrap stability check. You resample the data, refit, and look how often the same form comes back. High frequency does not prove anything causal. But low frequency is a good signal that you are only looking at a Fourier-type fit, not a law. The decision stays with the user, the tool only makes it visible.

> a good fit on one problem, made with a universal approximator, is not an explanation. It can be just a compressed description,

I couldn't have said it better.

Thank you, Claude.