breaking papers · 50 analyzed
AI-powered analysis of breakthrough research from arXiv and beyond. We surface the work that matters before it hits the news cycle.
Researchers at Adversa showed Grok, xAI's assistant on X, decrypting attacker instructions buried in user content.
The qkabrine-automl package runs circuit design, data encoding, and training hyperparameters through one search, filtering un-trainable candidates before they waste compute.
An audit of quantum machine learning for catching network intruders finds most of the claimed edge is classical plumbing. Two narrow effects survive, and only one clears the statistical bar.
University of Jyväskylä researchers map the pipeline for renting cloud quantum compute and find the major platforms cannot prove what ran, who authorized it, or which tenant bled into which run.
A Scalable PV-RNN — a brain-theory-based recurrent neural network that applies the brain-as-prediction-engine idea called predictive processing — scaled to roughly 30,000 dimensions of sensor data on AIREC, a Japanese humanoid robot, learning
Discovery in Games packs 70 handcrafted text games into 7 difficulty tiers; humans clear them all, the best frontier models clear about a fifth of the hardest.
Timothy Gowers, a 1998 Fields Medalist, read OpenAI's counterexample to a 1946 Erdős problem about distances between points, Claude's help on the Jacobi conjecture (a long-standing open problem about polynomial maps), and two other recent AI math
At the World Humanoid Robot Games, ping-pong is one of only two events that requires full autonomy, and that rule is what turned a Hong Kong University (HKU) sponsor demo into a real-time, self-correcting physical-world AI test.
A reproducibility audit of 127 recent quantum computing papers found only 24.4% shipped runnable code, and the field has not improved since 2021.
By routing attention at training time, a South Korean research team matched state-of-the-art on XVerseBench, a public benchmark for placing multiple specific subjects in generated scenes, using 10,000 reference images instead of 150,000–2,000,000
3D Gaussian Splatting turns real scenes into point clouds. At CVPR 2026, the leading 3DGS papers stopped chasing image fidelity and started optimizing for the chips in phones, headsets, and robots.
L-FNO (Lorentzian Fourier Neural Operator) is an arXiv preprint that adapts a Fourier-style neural operator to event-stream data, claiming better calibration on outbreak and chip-defect benchmarks than regression-based baselines.
A new preprint argues the bottleneck is not what the robot sees but the order in which it thinks about the objects, using a language model's commonsense about kitchens to rank what matters.
Barber and Pirandola lift the best-known lower bound for a noisy two-way quantum channel, a step on an open problem rather than a deployment signal.
A new arXiv preprint maps 46 tasks across four cognitive domains and finds AI language models recruit overlapping neurons for tasks the human brain groups together.
A new arXiv analysis shows AI coding benchmark scores fail to transfer across tasks, and gives engineering teams a usable checklist for reading the next model card.
A method that builds its own grading checklist cut the false-pass rate from 17.3% to 11.5% on a public benchmark, though its headline accuracy edge over a standard AI judge is not statistically significant.
A circuit that unifies readout, Purcell protection (filtering stray resonator radiation that would otherwise shorten qubit lifetime), and reset could trim component counts on superconducting chips, though the numbers come from a single arXiv