I’m not convinced GPU acceleration is a meaningful advantage for most charting use cases. Most dashboards don’t render enough data for it to matter. Once a chart is dense enough for rendering to become the bottleneck, it normally is already be too crowded to be meaningful.
Zooming can justify supporting larger datasets, but sampling/viewport culling and level of detail often avoid drawing unnecessary points...
I think meaningful is in the eye of the beholder. The library is designed such that the trace buffers are directly used as inputs to the WebGL2 drawing contexts to avoid unnecessary copying throughout the stack, which does make a difference when rendering on mobile and embedded devices with limited CPU but often having GPU resources available.
The good reason for worry about it is the same for data grid, list, scrolls and any other UI component that loads arbitrary data.
All UI, honestly, is only meaningful in what the screen size and our vision permit. END.
UNFORTUNATELY, you can't avoid that a user is writing "a___" and the source data has millions of things that start with `a` and all the others are dozens.
So, you can end with a massive influx of data, and sure the user see that big mess and wanna dial in, but in the meantime is nice if the UI not die in the process.
It depends on how much data you are planning on showing, but as you can see from the benchmarks its also more performant than other python charting libs for small data.
We also built this library for extreme customization with CSS/Tailwind support so rendering large amounts of data is an important but not the only advantage.
Why settle for sampling when you can have the whole dataset?
The spiral pattern is an excellent example. "Sure it looks like this when you zoom out, but when you zoom in, you can see the finer structure of the points..."
I’m not trying to justify whether I'd personally use it, I was simply raising the question because the tradeoffs are interesting, and the replies, including Evidlo’s, have already taught me something.
I tried the library before commenting, it’s a cool project. The performance improvements at larger scales I've found are real. I was mainly just trying to have a conversation to learn where we all learn!
“Just move on” seems like a curious response on a discussion forum, though :)
Oh, no worries. I thought you were another one of those naysayers, of which this site has far too many already. Hehe.
Btw, English isn't my first language, so I still struggle with it sometimes. Could you point me to the part of your comment where you asked that question? I can't seem to find it. Thanks!
It was not a explicit question where its easy to point at like a question mark (?) but I'd find any comment on a forum like environment to be trying to contribute to a discussion as a whole, where we can all share thoughts and counter thoughts constructively.
Anywhoo, XY seems like a cool lib. If you can find the usecase where you actually can leverage the power you should! Good job on Reflex.
Interesting; how do the examples compare to datashader?
Edit: for my use cases, I use napari (~1e7-8 points) if I need true interactivity; otherwise, datashader/holoviz, or even just fast-histogram's 2D histograms work.
For extremely large point clouds, these caveats[0] still apply. It irks me when people make dense scatterplots without any indication of just how dense some portions are.
Still, if it can indeed handle 1e10 points, that's pretty impressive.
I can imagine this useful to 'compress' gigabytes of data onto a 2d canvas quickly. For that, I appreciate the effort.
One thing that would be useful is to read up on Ed Tufte's principles of data visualization. Many graph libraries don't implement basic visualization principles to make they key point clear, easy to see while still keeping the full depth and complexity of data visible.
it's possible to render data out-of-core with XY, allowing it to render the entirety of OpenStreetMaps (that's 10,742,674,832 nodes!) with sub-second pan/zooms. it's a bit difficult to host online but you can try it out locally: https://github.com/reflex-dev/xy/tree/main/examples/osm
Check out mosaic from uwdata which works on top of Observable plot
Or plotly-resampler which works on top of plotly and uses the rust package tsdownsample to aggregate on the 4pixels per pixel shown level (to make antialias work)
the grammar of graphics approach really is a great abstraction, and I'd love to see xy work in that direction
Thanks! XY already uses a similar pixel-aware approach
long line and area traces are reduced in Rust using M4 to produce viewportsized extrema, that is then refined as you zoom. Dense scatter plots use a fixed-size density grid plus a representative sample
Interesting approach to large scale visualisation. Moving reduction into Rust and sending screen bounded data to WebGL seems much more sensible than pushing millions of raw points into the browser. How does it perform with real time updates? I am assuming it is much more performant? Any plans for a prod deployment?
Yes, it’s significantly faster than existing Python charting libraries for real-time updates.
Instead of serializing and sending the full dataset as JSON, it sends compact typed binary buffers and only the screen-relevant data reducing payload size and browser-side work.