
Contents
Snowflake Cortex AI is the set of AI services that run inside your Snowflake account. It covers SQL functions that call large language models, Cortex Analyst for questions over tables and Cortex Search for documents. Cortex Agents orchestrate them, and Snowflake Intelligence, now renamed Snowflake CoWork, is the chat app on top.
The appeal is simple. Your data stays under the same account, roles and policies you already run. The risk is also simple. It is easy to switch on and easy to spend money on. If the data underneath is not modelled properly, it gives confident wrong answers.
In the manufacturing and defence programmes I've worked on, the question was never "can Snowflake do AI?". It was "can we trust the answer, and can we prove who saw what?". This guide is written for that conversation.
Product names in this area have moved fast. Two changes matter. Snowflake announced Snowflake Intelligence as generally available on 4 November 2025, alongside Cortex Agents and the Snowflake-managed MCP server. Then at Summit in June 2026 it renamed the product Snowflake CoWork. Some SQL privileges and object names still use the old name.
Document AI is gone. Snowflake decommissioned the Document AI interface on 16 March 2026 and moved document extraction to the AI_EXTRACT function. If an SI proposal still mentions Document AI model builds, it is out of date.
Here is the stack as it stands in October 2026, with status and billing basis from Snowflake's documentation.
| Component | What it does | Status | Billed on |
|---|---|---|---|
| Cortex AI Functions | SQL functions such as AI_COMPLETE, AI_CLASSIFY, AI_FILTER, AI_AGG, AI_EXTRACT, AI_PARSE_DOCUMENT, AI_REDACT | GA, some functions in preview | AI Credits per million tokens, or per page for parsing |
| Cortex Analyst | Turns a business question into SQL over a semantic view | Available; since August 2026 Snowflake recommends calling it through Cortex Agents, and the REST API stays supported | Per message via the API, per token via Agents |
| Cortex Search | Hybrid keyword and vector search over text, used for retrieval | GA since October 2024 | AI Credits per GB of index per month, plus refresh warehouse |
| Cortex Agents | Plans a task and calls Analyst, Search, custom tools and MCP connectors | GA since November 2025 | Orchestration tokens plus each tool's own cost |
| Snowflake CoWork (formerly Snowflake Intelligence) | The chat app for business users, built on Cortex Agents | GA since November 2025; several features announced in June 2026 were labelled preview or "coming soon" | Tokens, plus tool and warehouse costs |
| Snowflake-managed MCP server | Exposes agents, Analyst, Search and SQL to outside AI clients | GA since November 2025 | Underlying tool costs |
The models come from Anthropic, OpenAI, Google, Meta, Mistral AI and others, plus Snowflake's own Arctic models for embedding and extraction. Which ones you can call depends on your region and your cross-region setting. More on that below, because it is the decision defence clients care about most.
Cortex Analyst does not read your tables and guess. It reads a semantic view: a schema-level object that names your business entities, measures, joins and synonyms. Only that metadata goes to the model to write SQL. The SQL then runs in your warehouse under the caller's role.
So the quality of every Snowflake CoWork answer about revenue, scrap rate or on-time delivery rests on that semantic view. In my experience this is where the real work sits. The hard part was agreeing metric definitions in the semantic layer, not data cleanliness. Defining "on-time delivery" once, in a way that sales, supply chain and finance all accept, takes more meetings than any model configuration.
Two features help. Verified queries let you store a known-good question and its SQL, so the model has worked examples. Snowflake also lets you run evaluations against them, which gives you an accuracy number instead of a demo impression.
The older approach, YAML semantic model files on a stage, still works. Avoid it for new work. Snowflake's own documentation recommends semantic views, and anyone with access to the stage can use a YAML model, which is a governance gap.
If your ERP data is not landed and modelled cleanly yet, start there. My guide on getting SAP, Oracle and Dynamics 365 data into Snowflake covers that step.
The strongest argument for Snowflake Cortex AI is governance. Cortex Analyst, Cortex Agents and CoWork respect your existing role-based access control, row access policies and masking. Model access itself is now controlled by role grants. The older account-wide model allowlist parameter is being retired during 2026.
There are nuances worth checking before anyone signs off.
- Cortex Search runs with owner's rights. Anyone with USAGE on a search service sees results from everything the owning role indexed. Build separate services per audience, not one big index.
- Agents use the caller's default role. If users carry broad default roles, the agent inherits that breadth. Tidy roles first.
- Cross-region inference moves prompts. The
CORTEX_ENABLED_CROSS_REGIONparameter controls whether inference can run outside your home region. Snowflake says data stays stored in the home region and prompts are not persisted elsewhere, but the payload does travel. Some newer organisations default toANY_REGION. The cross-region inference documentation recommends setting it explicitly if you have residency requirements. - The MCP server can bypass your agent. Snowflake advises exposing a Cortex Agent as the only tool for governed questions. Exposing raw SQL execution on the same server lets an outside client go around the agent's controls.
For defence, point 3 is the conversation. Setting the parameter to DISABLED or a regional value keeps inference within your boundary, but it narrows the models you can use and costs more per credit. Snowflake's US government regions keep requests inside the same compliance boundary. Its MCP documentation also states that a formal FedRAMP assessment of the MCP server has not been completed. Classified data has no place in any of this. Keep the pilot on unclassified, non-export-controlled data, and get your export control lead to agree the boundary in writing.
My AI governance framework guide covers the policy side.
AI features are billed in AI Credits, which are separate from the Platform Credits that pay for warehouses. Per Snowflake's AI pricing page, an AI Credit costs $2.00 on demand with global routing and $2.20 when you restrict inference to your region or disable cross-region routing. AI Credit pricing is the same across editions. Your Platform Credit capacity discount does not apply to it.
The rates below are list prices from the Snowflake Service Consumption Table effective 9 October 2026, checked October 2026. Enterprise agreements, AI Credit capacity tiers and new model releases will change the numbers.
| Item | List rate |
|---|---|
| AI Credit, global routing / regional routing | $2.00 / $2.20 |
AI_COMPLETE with llama3.3-70b (input / output) | 0.432 / 0.432 AI Credits per million tokens |
AI_COMPLETE with openai-gpt-4.1 (input / output) | 1.20 / 4.80 AI Credits per million tokens |
AI_COMPLETE with claude-sonnet-4-5 (input / output) | 1.80 / 9.00 AI Credits per million tokens |
AI_CLASSIFY or AI_FILTER | 1.62 AI Credits per million tokens |
AI_EXTRACT (arctic-extract) | 5.55 AI Credits per million tokens |
AI_PARSE_DOCUMENT, layout mode | 3.66 AI Credits per 1,000 pages |
| Cortex Search serving | 6.3 AI Credits per GB per month of indexed data |
| Cortex Analyst via the REST API | 67 Platform Credits per 1,000 messages |
The spread between models is large. On output tokens, claude-sonnet-4-5 costs about 20 times llama3.3-70b. Running a classification over ten years of maintenance notes on the wrong model is how AI budgets get blown in a week.
Three cost traps catch people. Multi-turn conversations reprocess the full history on each turn, so long chats cost more per question. Agent costs stack: orchestration tokens, then each tool, then the warehouse running the SQL. Provisioned throughput, if you reserve it, is non-cancellable and billed whether you use it or not.
The controls exist. Use resource budgets, per-user quotas and orchestration budgets on agents. Monitor the usage views, such as CORTEX_FUNCTIONS_QUERY_USAGE_HISTORY, CORTEX_AGENT_USAGE_HISTORY and SNOWFLAKE_COWORK_USAGE_HISTORY. Test every batch job on a sample first. My Snowflake pricing and credits guide covers the platform side.
The model is the cheapest part to change. The semantic layer is where the answers are won or lost.
Snowflake's own documentation is honest here. Cortex Analyst answers questions that SQL can answer. It does not do open-ended "why did margin fall?" reasoning, and it cannot reuse results from earlier queries. Snowflake also warns that agent responses and citations are not guaranteed to be accurate.
My judgment: treat every AI answer that feeds a decision as a draft that a named person signs. That matters more in defence, where a wrong figure can end up in a contractual or audit document.
Lock-in is the other risk. Semantic views, agents and search services are Snowflake objects. Your prompts and evaluation sets are portable; the plumbing is not. If you are still choosing a platform, my Snowflake vs Databricks comparison sets out that trade-off.
The first use case I'd pick in both sectors is the same shape: one structured question set and one document set, for one team.
In manufacturing, I'd make the first pilot a maintenance and quality assistant. A semantic view sits over work orders, downtime, scrap and non-conformance data from the ERP and MES. A Cortex Search service sits over maintenance procedures and closed quality reports. A maintenance planner asks "which lines had the most unplanned downtime last quarter, and what did we find last time?" and gets figures plus the relevant reports.
In defence, the same pattern works for supplier quality on an unclassified programme: delivery and quality metrics from ERP, plus supplier corrective action documents. Keep classified and export-controlled material out of scope entirely.
A 90-day Cortex AI pilot
Scope
Pick one team, 20 to 30 real questions and one document set. Set the cross-region parameter and roles before anything else.
Model
Build the semantic view and write verified queries for the top questions. Agree metric definitions with the business owner.
Build
Create the Cortex Search service per audience and one Cortex Agent. Set budgets and per-user quotas on day one.
Test
Run evaluations against the question set. Users rate answers. Compare with how long the same answers take today.
Decide
Scale, fix the data, or stop. Write down which, and why.
Measure four things, and record a baseline before you start.
- Answer accuracy against the verified question set, scored by the business owner.
- Time to answer compared with the current route, usually an analyst request or a report hunt.
- Cost per answered question, from the usage views, including warehouse time.
- Access exceptions: any case where a user saw data their role should not reach. The target is zero.
Before you buy anything, pull your current CORTEX_ENABLED_CROSS_REGION setting, list who has broad default roles, and pick the 20 questions your operations director asks every month. That one-page brief will tell you more about readiness than any vendor demo.
If you want a second opinion on the pilot scope, the residency settings or an SI's proposal, I help organisations with these decisions. You can get in touch here. For the Databricks equivalent, see my piece on Databricks Genie and Agent Bricks.
What is Snowflake Cortex AI?
Snowflake Cortex AI is Snowflake's family of AI services that run inside your Snowflake account. It includes AI SQL functions that call large language models, Cortex Analyst for questions over tables and Cortex Search for documents. Cortex Agents orchestrate the tools, and the Snowflake CoWork app serves business users.
Is Snowflake Intelligence the same as Snowflake CoWork?
Yes. Snowflake Intelligence became generally available on 4 November 2025 and was renamed Snowflake CoWork at Snowflake Summit in June 2026. Some SQL privileges and default object names still use the Snowflake Intelligence name.
Does Snowflake Cortex send my data outside Snowflake?
Models run inside Snowflake's service perimeter, and your stored data stays in your account's home region. With cross-region inference enabled, prompts and responses can be processed in another region and are not persisted there. If you have residency rules, set the CORTEX_ENABLED_CROSS_REGION parameter explicitly.
How much does Snowflake Cortex AI cost?
Most Cortex features bill in AI Credits, at $2.00 each with global routing or $2.20 with regional routing, list price checked October 2026. Rates vary by model and feature, from fractions of a credit to several credits per million tokens. Cortex Analyst via the REST API bills 67 Platform Credits per 1,000 messages. Enterprise agreements change these numbers.
What are Snowflake semantic views?
Semantic views are schema-level Snowflake objects that describe business entities, measures, relationships and synonyms. Cortex Analyst and Cortex Agents use them to turn questions into accurate SQL. They replace the older YAML semantic model files and are the main factor in answer quality.
What replaced Snowflake Document AI?
The AI_EXTRACT function. Snowflake decommissioned the Document AI interface and its PREDICT method on 16 March 2026. AI_EXTRACT pulls entities, lists and tables from documents in a single SQL call and bills on tokens.
Next step
Running an ERP programme right now?
If this article touched on a programme you are live in right now, a 30-minute conversation usually gets further than another week of internal analysis.




