Microsoft’s Skala Turns DFT Accuracy Into a Software Distribution Problem
Microsoft Research's Skala 1.1 pushes its deep-learning DFT model closer to hybrid-level accuracy at meta-GGA cost, but the bigger move is getting it inside the quantum chemistry software researchers...
Microsoft Research has spent the past ten months chasing one of computational chemistry’s oldest headaches: getting fast, cheap simulations to also be accurate. On August 20, the company announced Skala 1.1, an update to its deep-learning model for density functional theory (DFT) calculations. The more interesting part of the announcement is not the accuracy gain. It is that Microsoft is now pushing the model into the quantum chemistry software packages researchers already run every day, instead of asking them to adopt something new.
Table Of Content
What DFT Actually Computes
DFT is the workhorse method computational chemists and materials scientists use to predict how molecules and materials will behave without physically synthesizing and testing them in a lab. It works by approximating how electrons, which hold atoms together in chemical bonds, are distributed, instead of solving the full many-electron Schrodinger equation directly, which is too computationally expensive for all but the smallest systems. Predictive DFT could, in principle, replace some trial-and-error laboratory testing in drug discovery, materials screening, battery design and catalysis with simulation.
The catch is that the exact mathematical expression for the part of DFT that matters most, called the exchange-correlation functional, is unknown. Chemists have spent decades hand-designing hundreds of competing approximations instead, organized in what researchers call Jacob’s Ladder: each rung up adds more descriptors of the electron density to improve accuracy, at the cost of much heavier computation. Even with that effort, Microsoft says most current functionals still carry errors 3 to 30 times larger than the roughly 1 kcal/mol of “chemical accuracy” needed for a prediction to be useful in place of a real experiment.
Replacing a Hand-Designed Approximation With a Learned One
Skala, which Microsoft first introduced in October 2025, takes a different approach: rather than hand-designing the next rung of the ladder, it trains a neural network to learn the functional directly from high-accuracy reference calculations. At launch, Microsoft reported that Skala reached close to chemical accuracy on atomization energies, with an error of 0.85 kcal/mol on a single-reference test set, while costing only about 10% of the compute of standard hybrid functionals and about 1% of the cost of the more expensive local hybrids it was built to compete with.
The model itself is small by modern AI standards. According to its Hugging Face model card, Skala 1.1 has about 385,000 trainable parameters, a tiny footprint next to a large language model, trained on roughly 78,000 reactions with coupled-cluster-derived atomization energies from Microsoft’s MSR-ACC/TAE25 dataset, plus additional data covering atomic properties, transition metals and noncovalent interactions.
What’s New in Skala 1.1
Today’s release is framed as proof that Skala keeps improving without a full redesign. Trained on 2.5 times more data than the first public version, Skala 1.1 now ranks first, earning what Microsoft calls “gold medals,” in 32 of the 55 categories of GMTKN55, a peer-reviewed benchmark database of 1,505 relative energies spanning thermochemistry, reaction barriers and noncovalent interactions that has become a standard yardstick for judging density functionals since it was published in 2017. Microsoft says that at the computational cost of a meta-GGA functional, one of the cheaper rungs on Jacob’s Ladder, Skala 1.1 now outperforms the best, most expensive global hybrid functionals across those categories.
Alongside the model update, Microsoft is publishing what it calls a “living benchmark”: a continuously updated, public performance record that will track every future Skala release, rather than a one-time set of launch-day claims. That is a real departure from how most model benchmarks get reported, as a snapshot that rarely gets revisited once the announcement cycle moves on.
The Real News Is Distribution, Not the Model
The headline on Microsoft’s own post, “broadening access to Skala,” points at the actual strategic bet. Skala is available today inside CP2K, a widely used open-source simulation package for atomistic modeling, through GauXC, a third-party library for evaluating exchange-correlation functionals that Microsoft extended with support for PyTorch-based models like Skala. Microsoft says it is also being integrated into four more packages: Psi4, FHI-aims, ORCA and VASP, covering a large share of the software researchers in computational chemistry, materials science and catalysis already rely on. That distinction matters: the CP2K integration is finished and shipping now, while the other four are still in progress.
Few working scientists will switch software packages just to try a new functional, no matter how good its benchmark numbers look. A functional that exists only as a standalone research release competes for attention with a lab’s existing, validated workflow. By lining up integrations with five of the field’s established packages instead of shipping Skala as a separate tool, Microsoft is betting that adoption depends less on pushing the accuracy ceiling higher and more on how little friction it takes for a working scientist to try it inside software they already trust.
Open Weights, With a Caveat
Skala’s code and model weights are both released under the MIT license, and the model installs directly with pip install skala, with GPU-accelerated builds available for CUDA 11, 12 and 13, according to its GitHub repository. That is a more open release than most commercial AI models, which are typically gated behind a paid API.
It comes with a real caveat, though. Skala’s own model card describes it as “not a production model” and states that users “are expected to have a basic understanding of the field of quantum chemistry and density functional theory.” This is not a consumer tool. It is aimed squarely at researchers who already know how to run a DFT calculation and are deciding which functional to plug in.
Why It Matters Beyond Chemistry
Skala fits a pattern showing up across AI-for-science efforts: instead of building ever-larger general-purpose models, labs are increasingly training small, specialized networks (Skala’s roughly 385,000 parameters against the hundreds of billions in a frontier language model) to replace one well-defined, decades-old approximation. The pitch is that a model this narrow can be validated against known physics and a peer-reviewed benchmark in a way a general-purpose chatbot cannot. Whether that translates into scientists actually adopting it now comes down to something far more mundane than model architecture: whether it shows up, with one line of configuration, inside the software they already have open.








No Comment! Be the first one.