GPT-5.6 Sol and Codex automate routine calibration of superconducting qubits

On September 8, 2026 a case study described how GPT-5.6 Sol, connected to Codex and to laboratory control software, was used to run calibration workflows on superconducting qubit devices at MIT’s Engineering Quantum Systems Group (EQuS). The setup let the agent select measurement parameters, operate hardware, analyze returned data and decide subsequent steps in a sequence of interdependent measurements.

How the agents were used

The team tested the system on an uncalibrated six-qubit chip of a standard type the group uses for fabrication benchmarking. The chips sit inside dilution refrigerators and are controlled entirely through software: microwave pulses probe qubit transitions, digitized signals are analyzed, and results determine pulse settings and how long a qubit retains quantum information.

Using measurement-specific skills provided to Codex, GPT-5.6 Sol identified transition frequencies, calibrated control and readout pulses, and determined coherence-related timings when signals were clear. In those cases the agent completed standard calibration sequences with little researcher intervention and saved results for subsequent measurements.

Limitations and current practice

The team reported that GPT-5.6 Sol required more time and occasional researcher guidance when experimental signals were weak or noisy. Interpreting ambiguous physical results remained challenging, and experienced researchers could sometimes identify optimal settings faster than the agent.

Beatriz Yankelevich said allowing agents to run routine measurements reduced time spent monitoring calibrations and let her concentrate on data analysis, experiment design and planning. She reported being able to check progress remotely and steer agents if needed. For novel experiments, she assigns Codex agents narrower goals and relies on their ability to write, modify and test control and analysis code against real measurements.

The EQuS group now uses agents regularly for routine chip characterization, which can otherwise take researchers several days per device. The case study indicates agents handled clearly defined workflows autonomously but still depended on human oversight for ambiguous or noisy results.


Original source: OpenAI News

Leave a Comment