What Frontier AI Actually Costs, and What Each Provider Keeps: The September 2026 Price and Privacy Table

September 2026 frontier AI price and privacy table with per-provider output token costs

15 min read·Published Sep 29, 2026

0 0 votes
Article Rating

Part of Start Here

Six dollars. That is the published price, on 29 September 2026, of one million output tokens from Qwen3.8-2.4T-A95B, the largest model in Qwen’s 3.8 family, as listed on DeepInfra’s model page and on OpenRouter’s batch pricing page. Six dollars buys a million generated tokens from a 2.4 trillion parameter model that no consumer machine reviewed here can hold. The cheapest published rental that fits a frontier mixture of experts model in the research behind this page is about $18.32 per hour, from Spheron’s GPU recommender for GLM 5.3 Flash. The $6.00 is not a discount on a local setup. It is the price of not having one.

The privacy half of the same month is a different kind of number. On the same day, a Google support page stated: “When saved media is used to train our AI models, it is disconnected from your Google Account and retained for up to 4 years.” Another Google page adds that this holds “even if you delete the original activity.” Anthropic states that chats may be retained in a de-identified format “for up to 5 years” when model improvement is on. Microsoft states that it will soon keep search data from authenticated users “for up to 5 years.”

The price is public. The retention is public. The numbers are simply never printed next to each other, which is what this page does. It reports published prices and published policy, quotes the providers’ own documents, and labels what is missing, unverified or contested. It does not rank providers and does not recommend one.

The price table: hosted inference, per million tokens

Prices in this table were read on 29 September 2026. The first block comes from Venice‘s public pricing page, which publishes a privacy class for each model but no retention period. The last two rows come from DeepInfra’s model page for Qwen3.8-2.4T-A95B and OpenRouter’s batch pricing page. All prices are US dollars per million tokens.

Model Input per 1M Output per 1M Context window Privacy class as published
GLM 5.3 $1.75 $5.50 1,000K Listed Private. No retention period published on the price page
GLM 5.3 Flash $0.15 $0.50 1,049K Listed Private. No retention period published on the price page
GLM 5.3 Flash, end to end encrypted endpoint $0.16 $0.54 1,000K Listed E2EE, Private. No retention period published
DeepSeek V4 Pro $1.65 $3.30 1,000K Listed Private. No retention period published
DeepSeek V4 Flash $0.14 $0.28 1,000K Listed Private. No retention period published
DeepSeek V4.1 Flash $0.38 $1.50 1,000K Listed Private. No retention period published
GPT OSS 120B, end to end encrypted endpoint $0.13 $0.65 128K Listed E2EE, Private. No retention period published
Gemma 4 31B Instruct $0.12 $0.36 256K Listed Private. No retention period published
Qwen3.8-2.4T-A95B, DeepInfra $2.00 $6.00 262,144 Not stated on the page reviewed
Qwen3.8-2.4T-A95B, OpenRouter batch tier $2.00 $6.00 1,010,000 Not stated on the page reviewed

Two things stand out. The spread between the cheapest and the most expensive output price in this table is more than twenty fold, from $0.28 to $6.00 per million output tokens. And for most rows, the price page says nothing at all about how long anything is kept. The retention answers live in other documents, which is where the rest of this page goes.

The retention table: what each provider’s own documents state

This table summarizes the providers’ own published text. The quotes themselves are in the next section, with the document named.

Provider and product What its own documents state Document
Google Search Media used to train AI models is disconnected from the account and retained for up to 4 years, even if the original activity is deleted Search Services History help pages
Google Gemini, consumer Auto-delete default is 18 months, selectable to 3, 36 or indefinite. Chats reviewed by human reviewers are retained for up to three years Gemini Apps Privacy Hub
Google Gemini, Workspace With history off, new chats are saved in the account for up to 72 hours Google Workspace privacy hub
Microsoft Copilot, older app Conversation activity stored 18 months by default, with an opt out from training Copilot privacy FAQ
Microsoft Copilot, app updated 18 August 2026 Prompts, responses and file contents are not used to train foundation models Microsoft activity history page
Microsoft Bing 18 month personalization cut off, and up to 5 years for authenticated users to come Microsoft search history page
OpenAI API Not used for training by default, abuse monitoring logs up to 30 days, Zero Data Retention available to eligible customers API data guide and Zero Data Retention announcement
OpenAI ChatGPT, consumer Data sharing is on by default on personal workspaces, with an opt out OpenAI help page
Anthropic Claude, consumer Up to 5 years de-identified with model improvement on. Flagged material up to 2 years, trust and safety scores up to 7 years Anthropic consumer retention page
xAI Grok Deleted conversations and Private Chat deleted within 30 days SpaceXAI Privacy Policy
GitHub Copilot Free, Pro and Pro+ Interaction data used for training from 24 April 2026 unless opted out. Business and Enterprise accounts unaffected GitHub changelog, 25 March 2026

What each provider’s own documents say, in their own words

Google Search. The Search Services History help page states: “In the past, saving history and personalized recommendations were managed by Web & App Activity and Search Personalization settings. Going forward, you’ll be able to manage history and personalized recommendations independently through 2 new settings.” The covered services are Search, Maps, Shopping, Flights, Hotels, Translate and News. Saved history can include, per the same page, “Your search queries… Results you view on Google Search. Generative AI responses in AI Mode or Ask Maps. Recordings and transcripts from Search Live and voice searches. Images you upload in Google Lens, AI Mode, or Google’s Try-on tool.” The retention line appears in two places: “Data used to train our AI models is disconnected from your Google Account and retained for up to 4 years,” and “Saved media that has already been selected to train AI models is no longer connected to your account and is kept for up to 4 years, even if you delete the original activity.” The opt out has a stated limit: “If you turn this Save Media subsetting off, previously saved media isn’t deleted and may continue to be used to improve Google technologies unless you delete it from your account.”

Google Gemini. The Gemini Apps Privacy Hub, self dated 10 August 2026, states: “You can change your auto-delete setting in Gemini Apps Activity from the default of 18 months to 3 months, 36 months, or indefinite.” It also states: “Chats reviewed by human reviewers (and related data like your language, device type, location info, or feedback) are not deleted when you delete your activity. Instead, they are retained for up to three years.” On the Workspace side, Google states: “When Gemini conversation history is off, new chats are saved in user accounts for up to 72 hours so Google can provide the service and process any user feedback. This activity doesn’t appear in Gemini Apps Activity.”

Microsoft. Two Microsoft documents are live and they do not say the same thing, and no single Microsoft page reconciles them. The Copilot privacy FAQ, which the page itself says applies to the older Copilot app, states: “By default, we store conversation activity for 18 months. You can delete individual conversations or your entire conversation history at any time.” It also states: “You can opt out of use of your conversation activity for model training at any time.” The page for the updated app, available as of 18 August 2026, states: “Prompts, responses, and your file contents when using the Microsoft Copilot app aren’t used to train foundation models.” For Bing, Microsoft states: “After 18 months, we no longer use cookies and other identifiers that could be linked to a specific account or device to personalize your search experience. In the coming months, we will retain data from authenticated users for up to 5 years to help improve our products and services.” The same page states: “When you clear your search history, some copies may be retained within Microsoft’s backup systems for a limited time.”

OpenAI. The API data guide states: “As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us).” It also states: “By default, abuse monitoring logs are generated for all API feature usage and retained for up to 30 days, unless longer retention is required by law, or is reasonably necessary to protect our services or any third party from harm.” For consumer ChatGPT, OpenAI states: “If you are on a ChatGPT Plus, ChatGPT Pro or ChatGPT Free plan on a personal workspace, data sharing is enabled for you by default, however, you can opt out of using the data for training.” OpenAI’s Zero Data Retention announcement, dated 19 August 2026 with an update on 22 September 2026, states: “Zero Data Retention gives eligible API customers a clear promise: OpenAI does not retain their prompts or model responses after a request is processed.” The same page adds one exception: “Images flagged for potential CSAM will continue to be retained for manual review and reporting purposes, even in Zero Data Retention deployments, as they are today.”

Anthropic. The consumer retention page states: “If you allow us to use your chats or coding sessions to improve Claude, we may retain your data in a de-identified format for up to 5 years in our model training pipelines.” It also states: “We retain inputs and outputs for up to 2 years and trust and safety classification scores for up to 7 years if your chat or session is flagged by our automated trust and safety systems as violating our Usage Policy.” And: “Your Incognito chats are not used to improve Claude, even if you have enabled Model Improvement in your Privacy Settings.”

xAI. The policy page is titled “SpaceXAI Privacy Policy” and states an effective date of 24 August 2026. It states: “For example, when Private Chat is turned on, conversations will not appear in your conversation history and your conversations will be deleted from SpaceXAI systems within 30 days unless it is necessary that they be kept longer for legal, compliance, or safety purposes. Further, if you choose to delete any or all of your conversations or if you choose to delete your account, we will delete the data within 30 days unless it is necessary to retain the data for legal, compliance, or safety purposes.” On training sources, the same page states: “We use information that is publicly available on the internet to train our models and provide resulting Output. While we do not intentionally seek out personal information, we understand that there is personal information incidentally included in these datasets.” And on connected Google content: “For users who opt to connect to Google Apps via Google OAuth, SpaceXAI shall not use any Google Apps content for any of its internal AI or other training purposes (such as training its machine learning models), including developing new products or services based on such content.”

GitHub. GitHub announced that from 24 April 2026 it will use interaction data from Copilot Free, Pro and Pro+, stating: “from April 24 onward we will begin using interaction data, specifically inputs, outputs, code snippets, and associated context, from Copilot Free, Pro, and Pro+ users to train and improve our AI models unless you opt out.” It states the mechanism: “Not interested? Opt out in settings under ‘Privacy’. If you previously opted out of the setting allowing GitHub to collect this data for product improvements, your preference has been retained.” And it states a boundary that matters to developers: “This update does not change our treatment of private repository source code stored on GitHub. We do not use private repository content at rest to train AI models. The interaction data covered by this update (e.g., prompts, suggestions, and code snippets generated during your use of Copilot) may be generated while you are working in a private repository, but we are not accessing or training on the stored contents of that repository at rest.” Enterprise and organization accounts are excluded: “These changes apply only to individual consumer accounts. Enterprise and organization-provided accounts remain governed by your Data Protection Agreement. Your data will not be used for AI training.”

The arithmetic, labelled as arithmetic

The multiplications below are this page’s own arithmetic, not published figures.

A hundred million output tokens costs $600 on the Qwen 2.4T row and $28 on the DeepSeek V4 Flash row. A billion output tokens is $6,000 and $280. The twenty fold spread is not a rounding difference between vendors. It is the difference between one model that needs eight 96GB accelerators and one that runs on a mid range card you already own.

Compare that to the local route, with the prices published separately. A used RTX 3090 rents at $0.22 per hour on RunPod’s Community Cloud, per its pricing page. A four hour a day user spends $321 a year at that rate, which is under the $600 of output tokens in the paragraph above and comes with a card attached instead of a bill. The catch is the model. The 2.4 trillion parameter class does not run on any single card reviewed here, and GLM 5.3 needs roughly 459.5GB at Q4_K_M, which is the published third party estimate this page relies on and labels as third party.

What is missing, unverified, or contested

First, the hosted pricing page used for the first block of the price table publishes a privacy class for every model and no retention period for any of them. That is not a contradiction. It means the price page is not the document that answers the retention question, and treating a privacy class label as a retention promise would be a leap this page does not make.

Second, DeepInfra’s model page and OpenRouter’s batch page say nothing about retention for Qwen3.8-2.4T-A95B. Their prices are published. Their retention terms were not found in the pages reviewed.

Third, Microsoft has two live documents that conflict on Copilot retention, and the company has not published a single page reconciling them. Both are quoted above so you can see the conflict rather than take a summary of it.

Fourth, xAI’s privacy page is titled “SpaceXAI Privacy Policy” and refers to “SpaceXAI LLC”, while its own contact section still lists “x.AI LLC”. The page text is reported as it stands. Whether that reflects a completed corporate renaming is not established, and this page does not assert it.

Fifth, the $18.32 per hour rental figure comes from a GPU recommender page rather than a signed contract, so treat it as an estimate published by the vendor, not a price you can book at.

Sixth, the retention rows for consumer products describe what the documents state, not what happens in practice. No independent audit of any of these retention periods is cited here, because none was found.

Frequently asked questions

Which provider keeps your data the longest?

On the documents reviewed here, the longest single figure is Anthropic’s: if you allow model improvement, chats may be retained in a de-identified format for up to 5 years in its model training pipelines, and flagged material is kept for up to 2 years with trust and safety classification scores up to 7 years. Google states that media used to train its AI models is disconnected from your account and retained for up to 4 years, even if you delete the original activity. Microsoft states that it will retain data from authenticated users for up to 5 years to improve its products. This page does not rank providers. It reports each provider’s own text and lets you weigh it.

Does deleting a chat delete the data?

Not always. Google’s own help page states that saved media already selected for AI training is kept for up to 4 years even if the original activity is deleted, and that turning the setting off does not delete what was already saved. Gemini’s privacy hub states that chats reviewed by human reviewers are not deleted when you delete your activity and are retained for up to three years. Anthropic states that Incognito chats are not used to improve Claude even when model improvement is enabled, which is the cleanest opt out of the three. OpenAI states that abuse monitoring logs are retained for up to 30 days even when data is not used for training.

What is the cheapest and the most expensive row in the price table?

The cheapest output price in the table is DeepSeek V4 Flash at $0.28 per million output tokens, with input at $0.14. The most expensive is Qwen3.8-2.4T-A95B at $6.00 per million output tokens on both DeepInfra and OpenRouter’s batch tier, with input at $2.00. That is a spread of more than twenty fold. A hundred million output tokens is $600 on the expensive row and $28 on the cheapest one. Those two multiplications are this page’s own arithmetic, not published figures, and the model at the expensive end is the one that needs rented accelerators rather than a card in your case.

Sources

0 0 votes
Article Rating
Published
Categorized as Blog
The Thrifty Dev, author at thethriftydev.com

By TheThriftyDev

Building smart with AI and automation. No fluff, just results.

0 0 votes
Article Rating
Subscribe
Notify of
guest
0 Comments
Most Voted
Newest Oldest
TheThriftyDev Dispatch
Quit Google in One Weekend

The 48-hour migration playbook: what to move first, what to keep, and the exact apps that won't make you regret it on Monday.

No spam. Practical privacy, AI, backup and tool drops. Unsubscribe anytime.
0
Would love your thoughts, please comment.x
()
x