When buying an AI subscription, users should look at two things: how good the model is, and where they are allowed to use the inference they’re paying for.
There are two broad models. In the first, subscription inference stays inside first-party apps and agent harnesses. The challenge with this approach is that no lab can build a first-party product for every workflow or edge case. Labs therefore have to choose between building increasingly full-featured applications with SDKs, or keeping the core agent harness relatively minimal while making it highly extensible, closer to the Pi or DeepSeek-style approach. In either case, the subscription inference remains confined to environments the lab controls.
The second model is more portable: subscription inference can travel with the user. Labs could allow OAuth-based or headless API access so that users can consume the inference included in their subscription from third-party apps and agent harnesses. They could still build best-in-class first-party products with proprietary features available only through their own subscription experience. Users may subscribe because those products are valuable, while also gaining the freedom to use the same underlying inference elsewhere.
That changes what an AI subscription represents. Instead of paying for access to one app, users are effectively paying for access to inference that can work across many apps and agent harnesses.
Rate limits would still matter. As frustrating as five-hour rate limits can be, especially when they interrupt a task midway through, they help make generous and potentially subsidized usage sustainable at a fixed monthly price. Users who consistently need more capacity could move to higher subscription tiers or paid API usage.
This could also become a stronger acquisition model for AI labs: build the best first-party apps, but let users take their subscription inference with them.