
Tech • AI • Robotics • Game
Shopify is using large language models to build systems that do more than automate routine work, including a digital twin that predicts merchant outcomes and management tools that flag project risks before delays occur.
Shopify’s technology leadership describes businesses as sequences of actions, much like text is a sequence of tokens. In that framing, a company can be modeled through events such as launching products, adding payment methods, taking loans, changing shipping speeds, or starting ad campaigns. That sequence-based approach enables a machine-learning system to predict likely next actions and outcomes.
The company’s HST prediction system creates a digital twin of a business and tests interventions on the model rather than on the merchant directly. Teams can run counterfactuals such as offering a loan, accelerating shipping, or launching ads, then identify which sequence of actions is most likely to increase growth. Those recommendations can then be deployed in production through financing offers, prompts, or communications.
The central argument is that AI’s biggest value is not only raising the productivity floor by speeding up known tasks, but raising the capability ceiling by enabling work that previously would not have been attempted at all. That distinction matters because labor savings are relatively easy to count, while entirely new products or decision systems create value that would otherwise have been absent. The digital-twin project is presented as an example of a capability that was effectively out of reach before current models.
The same thinking applies beyond product development. A difficult mathematical problem involving a Wasserstein loss function and transforming distributions toward Gaussian form had remained unsolved for years, yet became tractable through iterative collaboration with a model. The result was neither pure automation nor a model acting alone, but a “centaur” workflow in which human judgment and model generation were jointly necessary.
Shopify says it has become highly precise at measuring AI’s effect on execution, going beyond simple counts of pull requests. Internal systems estimate project complexity and compare output after normalizing for that complexity, allowing teams to quantify how many more projects can be completed and how much faster. That makes floor-raising effects measurable even as ceiling-raising effects remain harder to express in a conventional ROI formula.
One major surprise was how useful models became for management rather than just coding. Internal tools now analyze organizational signals to identify projects likely to slip before deadlines are missed and to detect signs of dissatisfaction or friction inside teams. That shifts a technical leader’s role away from monitoring whether individual tasks were completed and toward supervising engineers who are themselves supervising model-driven workflows.
The operating principle is to start with the largest available model because that is where ceiling-raising gains are most likely to appear. Shopify has adopted effectively unlimited token access internally and invests heavily in inference, training, and optimization so staff can use the strongest models available. The view is that leaders cannot judge what they are missing unless they first test the highest-capability systems.
At the same time, the company advocates a barbell strategy. For coding, development, testing, and research, teams should spend on the largest models they can justify because that is where discovery and leverage are highest. For production services, they should tune for cost and throughput, often using cheaper or fine-tuned models, then redirect the savings back into high-end development usage.
Faster model-driven execution also changes governance needs. Because models are skilled at inferring what users seem to want, they can produce polished but misguided outputs if the initial prompt reflects the wrong objective. That makes technical judgment and careful supervision more important, especially when multiple AI systems or agents are working in parallel and leaders must prevent conflicts, runaway token usage, or low-quality decisions from shipping too quickly.
Shopify’s approach treats AI as a system for discovering actions and products that were previously infeasible, not merely for reducing effort on existing work. The competitive edge, in this view, comes from pairing top-tier models with strong human oversight to expand what companies can actually build and manage.
Ask a question