The frontier, on the record

THURSDAY · 10 SEPTEMBER 2026

Unitary

Quantum · AI · Robotics

Quantum Computing

A reverse-built benchmark exposes how much quantum compilers waste — and closing the gap pays off on hardware

Evaluating a quantum compiler has always lacked a ground truth: nobody knows what the optimal compiled circuit looks like, so "better than the other compiler" is the best claim available. QROB sidesteps this by construction — it starts from circuits that are already directly realizable on the hardware and runs the mapping backward, keeping the inverse path as a known-feasible reference. Measured against that reference across systems from 9 to 156 qubits, NISQ compilers pay up to 24.1x the reference SWAP cost and fault-tolerant lattice-surgery schedulers up to 7.0x the reference makespan. The references also work as training data: a router supervised on them outperforms Qiskit's SABRE on 84.8% of real application circuits. Most importantly, the gap is physical, not cosmetic — on three 156-qubit IBM Heron-r2 processors, QROB reference realizations survive mirror-circuit tests at 1.65x the median rate of full Qiskit O3 compilation.

arXiv quant-ph

Quantum Computing

Quantum LDPC codes break the orthogonality barrier, hitting a 10^-8 error floor at 4% noise

Classical LDPC design says: raise the girth of the Tanner graph while keeping the degree distribution regular, and you get both good belief-propagation decoding and large minimum distance. Quantum LDPC codes have not been able to follow that recipe, because the orthogonality constraints between the two parity-check matrices collapse the girth and structurally cap the minimum distance. This construction relaxes the bind by using permutation matrices with controlled commutativity and enforcing orthogonality only on the active part of the construction, preserving regular check matrices. The concrete demonstration is a girth-8, (3,12)-regular [[9216,4612,<=48]] code that, under BP decoding plus low-complexity post-processing, reaches a frame error rate of 10^-8 on a depolarizing channel at 4% error probability — a rate-1/2 code operating far above where surface codes need to sit.

Quantum (journal)

A tensor-network heuristic pushes classical verification into regimes quantum simulators had to themselves

Claims of quantum simulation advantage in two-dimensional spin systems rest on a moving boundary: the regime where classical tensor-network methods stop converging. This PRX Quantum paper moves that boundary, introducing a heuristic that speeds up convergence of matrix-product-state simulations of out-of-equilibrium 2D dynamics. The practical consequence is verification rather than competition — experiments whose results previously could not be independently checked classically now fall inside reach. The author list is notable in itself, drawn largely from the group that has spent years on the classical-simulation side of Google's quantum supremacy claims.

PRX Quantum

Loading real data into a trapped-ion machine — and certifying it in 1,000 shots

Getting classical data into a quantum computer and then proving the state you got is the state you wanted are both bottlenecks, and the verification half is usually the worse one: shadow-overlap certification works well for random states but its sample complexity explodes for the structured states that practical algorithms actually use. This work prepares a structured complex state encoding a digitized acoustic signal on Quantinuum's H2-1 and measures 0.929 hardware fidelity using resource-minimal circuits, with no fault-tolerant overhead and no idealized assumptions. The verification contribution is the sharper one — a pre-measurement basis change cuts the certification parameter tau by more than ten orders of magnitude for structured targets, tightening the theoretical guarantees and bringing the whole procedure down to 1,000 shots.

arXiv quant-ph

A generative model patches the weak spot in sample-based quantum diagonalization

Sample-based quantum diagonalization builds a determinant subspace from hardware samples and diagonalizes it classically, but in strongly correlated regimes the relevant determinant space outruns what finite-shot sampling can capture, and accuracy degrades. Q-WAVE seeds a beta-annealed variational autoencoder with determinants from SqDRIFT Krylov circuits plus CISD, lets it learn the wavefunction's support structure in a continuous latent space, and generates new dominant determinants unconstrained by any fixed excitation hierarchy. The resulting compact wavefunction beats what the raw hardware samples support: sub-millihartree accuracy against full CI for H2O and N2 dissociation, sub-millihartree against CCSD(T) on 52-qubit ethylene, and chemical accuracy on the notoriously hard 60-qubit Cr2 case after a final perturbative correction.

arXiv quant-ph

One compiler, four hardware modalities, and the resource bill for early quantum chemistry

Resource estimates for fault-tolerant algorithms are usually locked to one architecture-hardware pairing, making cross-platform comparison guesswork. This framework re-compiles a circuit into each platform's instruction set and fault-tolerant operations, reporting physical-qubit count, time-to-solution and classical processing time end to end, and adds a transversal active volume compilation architecture for platforms with long-range logical connectivity. Benchmarked on 2D Fermi-Hubbard simulation and eigenenergy estimation of trimethylenemethane, it puts an early fault-tolerant chemistry demonstration at roughly 10^4 physical qubits — with runtimes spanning three orders of magnitude, from 10^2 ms on photonics and superconducting to 10^5 ms on neutral atoms.

arXiv quant-ph

Cheaper logical fanout across distributed error-corrected processors

Logical fanout — a CNOT from one control to many targets sitting on remote nodes — is a basic primitive for distributed quantum computing and an expensive one, since implemented directly between encoded blocks it burns large amounts of non-local communication. This construction exploits the structure of encoded blocks and available transversal logical operations to build distributed fanout circuits that preserve the logical action while reducing non-local operations, developed generally and illustrated on Bivariate Bicycle code blocks with a full accounting of gate, entanglement, depth and ancilla requirements. A distributed global GCZ gate exploiting concurrency in the transversal fanouts is worked as an application.

arXiv quant-ph

Explicit circuits for putting polymer reaction kinetics on a fault-tolerant machine

Copolymer kinetics is a natural target for quantum simulation because the number of distinguishable species grows exponentially with chain length, but the rate matrices governing the kinetics are non-unitary and must be block-encoded before any of that helps. This work builds explicit circuits for two living-polymerization models: single-monomer polymerization, whose lower-bidiagonal rate matrix goes through a sparse oracle and a two-term LCU, and two-monomer copolymerization, where a bijective integer labeling of polymer species yields a structured sparse matrix encodable by a five-term LCU or a sparse oracle. Simulations reproduce classical time evolution across near-random and blocky microstructures, and the resource count is O(log N) qubits with gate counts reaching 10^4-10^5 at 10^3 system qubits.

arXiv quant-ph

A hard limit on what one round of quantum distributed computing can do

Lower bounds in distributed quantum computing have generally been proved in the non-signaling or bounded-dependence models, which are weaker abstractions than actual quantum algorithms. This result goes further: one-way one-round quantum LOCAL algorithms cannot 4-color directed cycles with high probability, even with unbounded local computation and unbounded quantum message length, and the proof exploits the structure of distributed quantum algorithms directly. The technique connects the two fields by identifying local collision probabilities with the weighted multiplicative energy of matrix-space decompositions, and the bound follows from a dimension-independent weighted stability theorem for a directed noncommutative version of Mantel's theorem.

arXiv quant-ph

Fujitsu builds a diamond tin-vacancy QPU prototype bonded to photonic ICs

Fujitsu says it has fabricated a prototype diamond-spin QPU that integrates tin-vacancy color centers into photonic integrated circuits, operating at -271.6C — warmer than superconducting qubits require. Tin-vacancy centers are chosen over the more familiar nitrogen-vacancy variety for lower sensitivity to electric-field noise and brighter photon emission; the scaling story rests on heterogeneous material bonding and nanometer-scale thinning of the diamond. The roadmap runs to a multi-module prototype in 2027 and a 1,000-logical-qubit machine by FY2035, with an eventual link to Fujitsu's superconducting hardware. Single-source vendor disclosure via trade press, with no independent characterization and no fidelity or coherence figures.

Quantum Computing Report

QKD

The attenuator itself is a side channel in chip-based QKD

Integrated QKD transmitters rely on variable optical attenuators to bring pulses down to the single-photon level, and this npj Quantum Information paper reports that those attenuators glow: VOA-induced luminescence carries information out of the chip and opens a side channel against chip-based quantum key distribution. It is a component-level attack surface rather than a protocol flaw — the kind that has repeatedly forced QKD implementations to be re-secured after the mathematics was settled.

npj Quantum Information

Post-Quantum Crypto

a16z rebuilds its zkVM on lattices, claiming post-quantum security and a 2-3x speedup

Zero-knowledge virtual machines let one party prove a computation ran correctly without revealing its inputs, and most deployed ones rest on elliptic-curve assumptions that a cryptographically relevant quantum computer would break. a16z Crypto has rebuilt its Jolt zkVM on lattice-based cryptography and released it as Lattice Jolt, claiming both post-quantum resistance and proofs two to three times faster than the scheme it replaces. The speed figure is the vendor's own and has not been independently reproduced, but unlike most post-quantum announcements this one ships as running open-source code.

The Quantum Insider

AI & ML

US agencies accuse Chinese AI firms of industrial-scale model distillation — and suggest quietly downgrading their access

A joint advisory from US cybersecurity and intelligence agencies alleges that six China-based AI companies have made "systematic extraction" of American frontier-model capabilities a core part of their development strategy, naming Claude, GPT, Gemini and Grok as targets of distillation attacks conducted at industrial scale. The recommended countermeasure is the more striking part of the document: agencies urge US providers to identify Chinese users and covertly switch them to weaker models rather than deny service outright. Both the accusation and the remedy are contested, and no technical evidence for the extraction claims has been published alongside the advisory.

The Hacker News · Ars Technica

Seven people, 164,269 verified trajectories, and an open-weight model in the CyberGym top ten

The bottleneck in training capable cyber agents is not model scale but the cost of executable environments, reliable multi-turn supervision and access to strong teacher models. This data-centric framework attacks all three with five systems covering reasoning-signature analysis, cheaper teacher sampling, capability repair after model merging and conversion of expert interventions into trainable reasoning, feeding a data engine of resettable coding, vulnerability, CTF, kernel-history, exploit, firmware and device-backed environments. Only execution-verified and evidence-audited trajectories are kept — 164,269 of them. The three resulting checkpoints gain 23.76% on average across CyberGym and 10.49% on pooled CTF suites; Feyospace-s1 posts 63.24% verified success and ranks 10th overall on the official leaderboard, first among comparably sized models.

arXiv cs.AI

LLM judges of explanations are mostly grading themselves

Simulatability measures an explanation by how well it helps someone predict a model's outputs, and since human evaluation is expensive the field has been swapping in LLM simulators. This paper replicates ConSim's rankings across the tested datasets, explanation families and simulators, then shows the protocol is measuring something else. When class names carry meaning, a simulator can just solve the classification task and score well without using the explanation at all; anonymizing the classes creates the opposite failure, rewarding explanations that leak the hidden label mapping — a leak the authors expose with a classes-as-concepts baseline. Both point to the same shortcut: predictions ride on task priors while explanations move the needle only slightly.

arXiv cs.LG

A code-first benchmark where later events revoke earlier ones — and six models score under 10%

A long context is often not a record but a process: later events revise or revoke earlier information, so answering requires identifying which records are still valid, applying updates in order and reconstructing state from history. Text-first synthesis pipelines cannot verify data of this kind because the transitions and answer logic stay implicit. EvolveScaler inverts the order — human-authored operational specifications define state transitions, validity, difficulty and executable answer logic, an LLM synthesizes a self-contained simulator per specification, and deterministic replay produces reference answers and atomic checklists alongside the rendered natural-language event history. Across 117 task prototypes and up to 1,200 events per instance, the strongest model manages 59.3% avg@5 on the hardest tier while six models fall below 10%; 6,000 training examples transfer, gaining 5.25 points on average across eight independently built out-of-distribution benchmarks.

arXiv cs.AI

Generative sampling breaks the timescale wall on rare molecular transitions

Rare conformational transitions are the events that matter most in biomolecular function and the ones molecular dynamics is worst at reaching, since brute-force sampling has to wait out the timescale and enhanced-sampling shortcuts demand collective variables chosen in advance. This Nature paper builds a generative path-sampling framework guided by the committor — the natural reaction coordinate — that reconstructs transition pathways without predefined collective variables, and recovers not just the paths but the thermodynamics and kinetics behind them, at a cost the authors describe as acceptable.

Nature — research

A virtual cell built on perturbation proteomics, aimed at drug discovery

Virtual cell models have mostly been trained on transcriptomic perturbation data, which is cheaper to gather but a step removed from the proteins that drugs actually act on. ProteinTalks is built instead from temporal protein-abundance measurements across systematically perturbed breast cancer cell lines, and Nature presents it as an operational tool rather than a demonstration — evaluated across a range of drug discovery tasks.

Nature — research

Who gets credit when an AI does the mathematics

OpenAI's claimed AI-assisted resolution of the Navier-Stokes problem — reported last week and still not independently verified — has moved into a second dispute, over attribution. Science reports on the fight now underway about who deserves credit for a result that cost millions in compute and involved human mathematicians at every stage of framing, checking and writing up. The mathematical claim itself remains unresolved pending verification; what is playing out in the meantime is the discipline's first serious argument about authorship norms for machine-generated proof.

Science — current

Robotics

LLMs designing soft robots start passing the physics check — 96.2% feasible, three built for real

LLMs can turn a high-level spec into a robot design, but designs for bodies that deform and interact — tendon-driven continuum robots, say — tend to be physically invalid, because language-based reasoning has no grip on the consequences of embodiment. AID-SR closes that loop, feeding simulator-observed physical states back to the LLM designer as structured feedback, combined with semantic critique, human feedback and iterative refinement. Across a 14-task benchmark spanning reaching, grasping, locomotion and manipulation, 96.2% of generated designs pass the simulation feasibility check, and 26.7% go on to complete their task after standard RL training. Three of the designs were fabricated and completed their tasks physically; code and resources are released.

arXiv cs.RO

An open stack turns world-action pretraining into a controlled experiment

World-action models inherit knowledge from video-generative priors and route it into control, but existing systems tie backbone, representation, architecture, information flow, inference and data together so tightly that nobody can say which choice is doing the work. OpenWAM factorizes the design space into composable modules with unified training, inference, deployment and evaluation, then runs controlled experiments on what to inherit, how world and action learning interact and how the synergy scales. Three principles come out: transfer needs a capable generative backbone and a compact information-rich latent space; world-action synergy needs dedicated action capacity, explicit world-to-action flow and synchronized joint denoising; and embodied pretraining mainly buys out-of-domain generalization. OpenWAM-alpha, pretrained on about 6,400 hours of egocentric human and robot data, is released with the full stack — though the headline evaluation claims are qualitative.

arXiv cs.RO

A common yardstick for robots learning from human video

Learning manipulation from human video has produced promising results that are almost impossible to compare, because methods differ in assumptions, hardware and environment setup. RoboReel bundles real-world human demonstration videos, simulated robot trajectories and evaluation environments across ten manipulation tasks, with four test suites probing robustness to visual distractors and the ability to finish long-horizon tasks. Running over seven state-of-the-art algorithms, including the authors' VLA-based variants, through the same pipeline yields a clear negative finding: long-horizon tasks and tasks with tight tolerances remain out of reach for current models.

arXiv cs.RO

Hard safety guarantees for sampling-based MPC, computed online

Sampling-based MPC is flexible enough to run on almost any robot and has historically offered no hard safety guarantees; the reachability-based planners that do offer them pay for it with expensive offline pre-computation that limits which systems they can handle. This work computes guaranteed reachable-set overapproximations online with a fast interval-based pipeline, matching state-of-the-art reachability planner performance without the pre-computation step and scaling to systems those methods cannot handle at all. In a racing simulation it removes over 99% of safety violations, and it drives a model racecar on real hardware without crashing.

arXiv cs.RO

Robots that talk by moving — a covert channel hidden in policy noise

Robots already move in rich, articulate ways, and this paper asks whether the motion itself can carry messages. The method encodes arbitrary short payloads — an agent's current intent, say — as noise injected into any pre-trained policy's actions, leaving policy performance intact while making the message recoverable from remote sensing like video or motion capture. The appeal is robustness rather than bandwidth: this physical channel needs no extra hardware, no direct link to the robot and no working wireless, so it complements standard radio rather than replacing it. On real robots at 50 Hz, four of them jointly recover an 8-bit message at an aggregate 0.67 bits per second.

arXiv cs.RO