Choosing the AI model in Copilot Studio looks like a small decision. It’s one dropdown on your agent’s Overview page. But that dropdown controls how well your agent reasons, how fast it replies, how many Copilot Credits it uses, and even where your data is processed.
I learned this firsthand last quarter. A logistics client asked me to build two agents in the same week:
- A Contracts Q&A Agent for legal and procurement. It had to answer questions like “Can we end this contract early if the vendor misses two SLAs?”
- A Password Help Agent for IT that walks employees through a self-service password reset.
Same company, same environment, completely different needs. The contracts agent needed careful, multi-step reasoning. The password agent needed to be fast and inexpensive, because hundreds of employees use it every Monday morning. They ended up on different models, and I’ll use both agents as examples throughout this guide.
In this guide, you’ll learn:
- Where to change the primary AI model in Copilot Studio
- What the General, Auto, and Deep categories and the release tags mean
- Which admin settings control model access and data residency
- How your model choice affects Copilot Credit costs
- How to add deep reasoning or a different prompt model without switching the whole agent
- How to test models side by side before you publish
Model names, settings, and rates in this guide were checked against Microsoft’s documentation as of September 2026.
Why the AI Model Matters in Copilot Studio
Every agent in Copilot Studio has a large language model (LLM) behind it. With generative orchestration turned on, this model decides which topic, tool, or knowledge source to use, fills in the inputs, and writes the reply. If you’re new to building agents, start with my guide on how to create an agent in Microsoft Copilot Studio and come back here.
So when you swap the model, you change the brain that makes those decisions. A new model might follow your instructions more closely. It might also format dates differently or call a tool too eagerly. That’s why I treat a model change like a design change: it gets tested and signed off, just like a new topic.
One quick note. Copilot Studio now runs agents on different harnesses (the runtime layer between your agent design and the model). This guide covers agents on the standard harness, which is where most topic-based and knowledge-based agents run today.
Quick Answer: Which Copilot Studio Model Should You Use?
If you just need a starting point, use this table. The rest of the guide explains the trade-offs.
| If your agent… | Start with | Why |
|---|---|---|
| Answers high-volume FAQs or runs simple actions | A General model, such as the default GPT-5.5 Chat | Fastest replies and lowest cost |
| Handles an unpredictable mix of simple and hard questions | An Auto model (GPT-5 Auto is in preview, so test it first) | Adjusts effort per turn |
| Analyzes contracts or policies, or troubleshoots across systems | A Deep model, such as Claude Opus 4.7, or a General model with deep reasoning on specific steps | Better multi-step reasoning and citations |
| Runs a simple summary or extraction prompt | A Mini prompt model, such as GPT-4.1 mini | Billed at the lowest (basic) rate |
| Must keep data in your geographic region | A GA model without the cross-geo tag in your region | Avoids cross-region processing |
For production agents, stick to generally available (GA) models.
How to Change the AI Model in Copilot Studio
The primary model is the one your agent uses for orchestration and responses. Here’s how to change it:
- Sign in to Copilot Studio and open your agent.
- Go to the agent’s Overview page.
- Find the Model section.
- Open the dropdown and select the model you want.
- Ask a few real questions in the Test pane to check the new behavior.
- When you’re happy with the results, Publish the agent.

The Test pane uses the new model right away, but your users only get it after you publish. You can switch models at any time.
Each model in the dropdown shows small tags. One tells you what the model is good at (its category). The other tells you how mature it is (its release type). Reading these tags correctly is most of the job.
Copilot Studio Model Categories: General, Auto, and Deep
Copilot Studio groups models by what they’re built for. Here’s how I explain it to clients:
| Category | Built for | Speed | Cost | Good fit |
|---|---|---|---|---|
| General | Everyday chat and light grounding | Fastest | Lowest | FAQs, summaries, drafting, translation, simple actions |
| Auto | Mixed workloads; routes each turn dynamically | Varies | Varies | Helpdesk and employee agents with unpredictable questions |
| Deep | Deliberate, multi-step reasoning with tools | Slowest | Highest | Contract and policy analysis, long documents, multi-system troubleshooting |
General Models: Fast and Low-Cost
A General model is your workhorse. It answers FAQ-style questions, rewrites and summarizes text, and triggers simple actions quickly and at the lowest cost.
For the Password Help Agent, General was the obvious choice. The questions are short, the answers come from one SharePoint page, and the agent mostly calls a flow. A reasoning model would only add waiting time.
Auto Models: Effort That Adjusts per Turn
An Auto model decides how much effort each turn needs. A simple “hi” gets a quick reply, while a harder question gets more thought. It suits broad employee agents with a messy mix of questions. The trade-off is that latency and cost become less predictable.
Deep Models: Multi-Step Reasoning
A Deep model (often called a reasoning model) works through a problem before it answers. It breaks the task into steps and handles long documents with citations.
For the Contracts Q&A Agent, questions like “Which carriers allow price changes with less than 30 days’ notice?” require comparing clauses across several agreements. The Deep model was noticeably slower, but its answers were more complete and better cited.
Model Tags Explained: Default, GA, Preview, Experimental, and Retired
The second set of tags tells you how ready a model is for production. This is where makers most often get into trouble.
- Default: The model new agents start with, usually the best-performing GA model. Microsoft upgrades it over time, and it’s the fallback if your selected model becomes unavailable.
- Generally available (GA): Models with no release tag. Tested and suitable for production, though some have regional limits.
- Preview: Likely to become GA later, but not recommended for production yet.
- Experimental: For trying things out only. Expect variable quality, latency, and possible timeouts.
- Early access environment: Experimental models that appear only in environments on Microsoft’s early release cycle.
- Retired: A previous default model after an upgrade. You can keep using it for up to one month.
- Cross-geo: Your data might be processed and stored outside your organization’s geographic region.
One billing warning: if you publish an agent that uses a preview or experimental model, usage is billed at normal rates. “Preview” doesn’t mean “free.”
Copilot Studio Models Available in September 2026
Here’s what Microsoft documents for standard-harness agents in US commercial environments:
| Model | Provider | Category | Status |
|---|---|---|---|
| GPT-5.5 Chat | OpenAI | General | Default (GA) |
| GPT-4.1 | OpenAI | General | GA |
| GPT-5 Chat | OpenAI | General | GA |
| Claude Sonnet 4.6 | Anthropic | General | GA |
| Claude Opus 4.6 | Anthropic | Deep | GA |
| Claude Opus 4.7 | Anthropic | Deep | GA |
| GPT-5 Reasoning | OpenAI | Deep | Preview |
| GPT-5 Auto | OpenAI | Auto | Preview |
| Mistral Medium 3.5 | Mistral | General | Experimental (cross-geo) |
| GPT-5.3 Chat, GPT-5.4 Reasoning, GPT-5.5 Reasoning | OpenAI | General / Deep | Experimental, early access only (US) |
| Grok 4.1 Fast (Non-reasoning) | xAI | General | Experimental, early access only (US) |
| GPT-4o, Claude Sonnet 4.5 | OpenAI, Anthropic | General | Retired |
Government clouds (GCC, GCC High, and DoD) still use GPT-4o as the default.
This list changes often, so trust the tags in your own dropdown over any blog post, including this one.
Admin Settings, Data Residency, and External Models
This is the part that gets projects stuck in security review, so check it before you start testing.
Cross-Geo Models and Data Residency
A GA (cross-geo) tag means the model is production-ready, but your data might be processed outside your region. Which models carry the tag depends on where your environment lives:
- Anthropic Claude models are GA in the United States but cross-geo in Asia, Europe, and the United Kingdom.
- GPT-4.1 and GPT-5 Chat are GA in the United States, Europe, and the UK, and cross-geo in Asia.
- Mistral Medium 3.5 is experimental and cross-geo in all regions.
For a European client with strict residency rules, the cross-geo tag alone can rule out a model. I check Microsoft’s region table before I test anything.
How to Turn On Preview and Experimental Models
If a maker says “I can’t see GPT-5 Reasoning,” an environment setting is usually off. In the Power Platform admin center:
- Turn on Preview and experimental AI models for the environment. Makers can’t see preview or experimental models without it.
- For experimental models, also turn on Move data across regions. This is managed at the tenant level and allows data to be processed outside your region.

How to Enable Claude, Mistral, and Grok Models
External models from Anthropic, Mistral, and xAI appear in the same Model dropdown, grouped under the provider name. An admin has to complete two steps first:
- Turn on external models for the environment (or environment group) in the Power Platform admin center.
- Allow each model provider in the Microsoft 365 admin center, for example Connect to Anthropic LLM, Connect to Mistral AI Models, or Connect to xAI.
These settings are separate from the preview settings, so turning on one doesn’t turn on the other. If an admin later removes access to an external model, agents using it switch to a suitable internal model, or show an error if none fits.
Microsoft also carries a strong warning for Grok 4.1 Fast (Non-reasoning). In Microsoft’s safety evaluations it scored lower on safety and jailbreak benchmarks and is more likely to produce harmful or explicit content. For anything customer-facing, I wouldn’t use it.
If you’re still getting your head around how environments, admin centers, and connectors fit together, my overview of what Power Platform is is a good primer.
How the AI Model Affects Copilot Credit Costs
Copilot Studio bills usage in Copilot Credits. The standard feature rates apply to every language model Copilot Studio provides. For example, a generative answer costs 2 Copilot Credits whether it comes from GPT-5.5 Chat or Claude Sonnet 4.6. Reasoning changes the math.
The Reasoning Premium
When an agent uses a reasoning-capable model, Copilot Studio bills two meters:
- The normal feature rate for what the agent did, such as 2 credits for a generative answer.
- Text and generative AI tools (premium) for the reasoning tokens, at 10 Copilot Credits per 1,000 tokens.
For the Password Help Agent, with around 600 conversations a week, that premium would add up fast. For the Contracts Q&A Agent, with maybe 40 questions a week from lawyers, it’s easy to justify.
So I always ask two questions: “How many conversations will this agent handle?” and “What does a wrong answer cost us?” High volume and low risk point to General. Low volume and high risk point to Deep.
When Microsoft 365 Copilot Licenses Cover the Cost
There’s an important exception. For employee-facing agents, usage is included (no Copilot Credit charge) when the user has a Microsoft 365 Copilot license and the agent runs under that user’s authenticated identity. If most of your employees are licensed, the credit math above mainly matters for unlicensed users, customer-facing agents, and autonomous agents.
To estimate costs before you build, try Microsoft’s Copilot Studio agent usage estimator.
How to Use Deep Reasoning Without Changing the Primary Model
There’s a middle path I use a lot. Instead of switching the whole agent to a Deep model, keep a fast primary model and turn on deep reasoning for specific steps.
- Make sure generative orchestration is on (Settings > Generative AI > Orchestration).
- In the agent’s Settings, turn on Deep reasoning (preview).
- In your agent instructions, add the word reason to each step that needs it.

The agent then uses the reasoning model where you asked, or where it decides deep thinking helps. Here’s the kind of instruction I used for the contracts agent:
1. Identify the supplier name and the question type (termination, pricing, liability, renewal).
2. Search the Contracts library for the matching agreement.
3. If the question compares two or more contracts, use reason to compare the relevant
clauses and explain the differences in plain language.
4. Always cite the clause number and the document name.
5. If you cannot find the clause, say so and suggest contacting the legal team.
A few things to know before you rely on it:
- Deep reasoning is still in preview and currently runs on the Azure OpenAI o3 model.
- It’s available in the United States and the EU (excluding the UK) and makes no data residency commitments.
- Every reason step slows the reply, and each one uses billable Copilot Credits. Use it on one or two steps, not every step.
- Check the Activity page afterward. The activity map shows a separate deep reasoning node wherever it ran, and you can expand it to see the steps and data used.
Deep reasoning is also handy for agents that run without a user, which I cover in my post on creating autonomous agents in Copilot Studio.
How to Choose a Different Model for a Single Prompt
This is my favorite cost saver. A prompt tool (built in the prompt builder) can use its own model, separate from the agent’s primary model.
The Password Help Agent has a prompt that turns the user’s issue into a clean ticket description before handing off to IT. That job doesn’t need the main model. A small, inexpensive model handles it fine.
Here’s how to change the model for a prompt:
- Open the prompt in the prompt builder (for example, from the agent’s Tools page).
- Select Model at the top of the prompt builder.
- Pick a model from the dropdown.
- Select the three dots (…) and then Settings to adjust Temperature and other options.
- Test the prompt with real sample inputs, then save.

The prompt model list includes a Mini category that you don’t see for the primary model. GPT-4.1 mini is the default prompt model. You’ll also find General options like GPT-4.1, GPT-5 chat, GPT-5.3 chat, and Claude Sonnet 4.6, plus Deep options like GPT-5 reasoning, GPT-5.2 reasoning, and Claude Opus 4.6.
Prompt Model Rates
Each prompt model is billed at a basic, standard, or premium rate:
| Prompt rate | Copilot Credits per 1,000 tokens | Copilot Credits per 10 responses | Typical use |
|---|---|---|---|
| Basic | 0.1 | 1 | Mini models for summaries and simple extraction |
| Standard | 1.5 | 15 | General models for richer writing and document work |
| Premium | 10 | 100 | Deep models for analysis and reasoning |
That’s a 100x spread between basic and premium, so using a Mini model for simple summaries is an easy win.
Temperature and Other Prompt Settings
- Temperature goes from 0 (predictable, the default) to 1 (creative). I keep it at 0 for anything factual.
- The temperature slider is disabled for the GPT-5 reasoning model.
- Microsoft still treats Anthropic Claude models in prompts as experimental, even though they don’t show a tag. They’re hosted outside Microsoft and aren’t recommended for production prompts yet.
- Prompts in Copilot Studio use Copilot Credits, while prompts in Power Apps or Power Automate use AI Builder credits first.
If you’ve used prompts in canvas apps before, the idea is the same. My walkthrough of the Power Apps prompt shows the Power Apps side, and my guide to extracting invoice details with AI Builder and Power Automate shows prompts inside a flow.
How to Test and Compare Models in Copilot Studio
Never switch models based on one lucky answer in the Test pane. Here’s the process I use.
Step 1: Build a Realistic Test Set
Collect 20 to 50 real questions, including the odd ones people actually type. For the contracts agent, two paralegals wrote down questions they get every week. For the password agent, I pulled phrases from old helpdesk tickets.
Save them as a CSV file with Question and Expected response columns, in that order:
Question,Expected response
"Can we cancel the Harbor Freight Lines contract early?","Yes. Clause 14.2 allows termination with 60 days' written notice if two SLA breaches occur in a quarter."
"Which carriers can raise prices with less than 30 days' notice?","Only Swift Coastal Transport, under clause 7.4 of its 2025 agreement."
"What is the liability cap in the BlueRoute agreement?","The cap is 12 months of fees, as stated in clause 11.1."
"i forgot my password and im locked out","Explain the self-service reset steps and offer to raise a ticket if the account is locked."
A test set can hold up to 100 test cases, and each question can be up to 1,000 characters.
Step 2: Run an Evaluation
- Open the agent’s Evaluation page.
- Select New evaluation, and then select Single responses.
- Choose Import to upload your CSV. You can also write questions yourself or generate a question set from your knowledge sources.
- Under Select test methods, add Compare meaning (checks whether the answer means the same thing as your expected response) and General quality. If your agent calls flows or tools, add Tool use as well.
- Select Run.

Step 3: Switch the Model and Run It Again
Change the primary model on the Overview page and run the same test set again. Comparing the two runs side by side makes regressions easy to spot. Results stay in Copilot Studio for 89 days, so export them to CSV if you need a record for sign-off.

Step 4: Look Beyond the Score
A score is a starting point. I also check:
- Latency: Lawyers will wait 15 seconds for a good answer. Someone locked out of their laptop won’t.
- Formatting: Did the model follow my rules for dates, lists, and citations?
- Tool use: Did it call the right flow or knowledge source, or did it guess?
- Refusals: Did it politely decline out-of-scope questions?
For the contracts agent, the General model missed cross-contract comparisons and the Deep model got them right. For the password agent, both scored about the same, so the faster, cheaper General model won.
Pro tip: Keep a “golden” test set for every agent. When Microsoft upgrades the default model, rerun the same set within a day. If something breaks, turn on Continue using retired models (covered next) and fix your instructions during the grace period instead of in a panic.
What Happens When Microsoft Upgrades the Default Model
When a new default model arrives, the old one gets the Retired tag, and agents on the default move to the new model automatically. If you need more time:
- Open the agent’s Settings page.
- In the Model section, turn on Continue using retired models.

You can then switch between the retired and upgraded models for 30 days after the upgrade. The preference also applies to future upgrades until you turn it off. I use this for regulated clients who must re-validate every change.
The Model Is Only Part of Answer Quality
When a client asks “Which model should we use?”, my short answer is: General for high-volume simple questions, Deep (or a single reason step) for analysis, GA only for production, and no cross-geo models if data must stay in your region.
But good knowledge sources and clear topics matter just as much. If your answers are weak, check your SharePoint list as a knowledge source setup, and consider adding knowledge files automatically with Power Automate so the model always has current content. For fixed, predictable steps like a reset request, a well-built custom topic in Copilot Studio or an agent flow is often more reliable than asking any model to improvise.
For bigger solutions, you can also split work across agents. A fast front-door agent can hand complex questions to a specialist agent running a Deep model. I explain the pattern in my guide to building a multi-agent solution in Copilot Studio. And if your team is still deciding between building in Copilot Studio or using what’s built into Microsoft 365, read my comparison of Microsoft 365 Copilot vs Copilot Studio first.
Best Practices for Choosing a Copilot Studio Model
- Treat model changes as design changes. Rerun your test set every time you switch models, even between versions from the same provider.
- Keep preview out of production. Preview and experimental models can time out, change behavior, or disappear, and published usage is still billed.
- Check data residency first. Cross-geo and experimental models might process data outside your region, which can fail a compliance review.
- Reserve reasoning for high-value questions. Reasoning models add a premium token charge on top of the normal feature rate.
- Use the cheapest model that passes your tests. A Mini prompt model or a single reason step often beats switching the whole agent.
- Know where the switches are. Missing models are usually blocked by the preview, data movement, or external model settings in the admin centers.
- Publish to apply. Test pane changes are instant, but users only get the new model after you publish.
Frequently Asked Questions
What is the default AI model in Copilot Studio?
As of September 2026, GPT-5.5 Chat is the default primary model for standard agents in commercial regions. Government clouds (GCC, GCC High and DoD) still use GPT-4o. Microsoft upgrades the default over time, so check the Default tag in your model dropdown.
How do I change the AI model in Copilot Studio?
Open your agent, go to the Overview page, and select a new model in the Model section dropdown. Test it in the Test pane, then publish the agent. Users only get the new model after you publish.
Which AI model is best for Copilot Studio agents?
There’s no single best model. Use a General model like GPT-5.5 Chat for high-volume FAQs and simple actions, and a Deep model like Claude Opus 4.7 for contract, policy or multi-step analysis. Test both with the same questions before deciding.
Can I use Claude models in Copilot Studio?
Yes. Claude Sonnet 4.6, Claude Opus 4.6 and Claude Opus 4.7 appear under Anthropic in the model picker. An admin must turn on external models in the Power Platform admin center and allow Anthropic in the Microsoft 365 admin center first. Outside the US, these models are marked cross-geo.
Why can’t I see preview or experimental models in the dropdown?
The environment’s Preview and experimental AI models setting is probably turned off. Experimental models also need the Move data across regions setting. An admin manages both in the Power Platform admin center.
Does changing the AI model change my Copilot Credit usage?
Standard feature rates, such as 2 credits for a generative answer, apply to every model Copilot Studio provides. Reasoning models add a premium charge of 10 Copilot Credits per 1,000 reasoning tokens. Prompt tools are billed at basic, standard or premium rates depending on the model.
What is deep reasoning in Copilot Studio?
Deep reasoning is a preview setting that lets an agent use a reasoning model for complex steps while keeping its faster primary model for everything else. Turn it on in the agent’s Settings, then add the word “reason” to the instruction steps that need it. It requires generative orchestration.
Can one Copilot Studio agent use more than one AI model?
Yes. The agent has one primary model, but each prompt tool can use its own model, such as GPT-4.1 mini for simple summaries. Deep reasoning can also send specific steps to a reasoning model.
What happens when Microsoft retires the default model?
Agents on the default move to the new model automatically, and the old model gets the Retired tag. Turn on Continue using retired models in the agent’s Settings to keep switching between the old and new models for 30 days.
How do I compare two models fairly in Copilot Studio?
Run the same test set of real questions on the Evaluation page, switch the primary model, and run it again. Compare the scores, then check latency, formatting, tool use and refusals before you choose.
Conclusion
Choosing the right AI model in Copilot Studio comes down to matching the model to the job. Use fast General models for high-volume questions, Deep models or a single reason step for careful analysis, and a Mini prompt model wherever it saves money. Test with real questions, respect your region and admin settings, and treat every model switch as a change that needs sign-off.
Start with one agent this week: build a 20-question test set, run it on your current model, and then try one alternative. The results will tell you more than any comparison chart, including the ones in this guide.
You May Also Like
- Create a SharePoint List Item Using Copilot Studio
- Add Event Triggers in Microsoft Copilot Studio
- Create a Custom Agent in Microsoft 365 Copilot
- Is Copilot Better Than ChatGPT?
- Build and Deploy a Smart HR Assistant Bot Using Copilot Studio

Hey! I’m Bijay Kumar, founder of SPGuides.com and a Microsoft Business Applications MVP (Power Automate, Power Apps). I launched this site in 2020 because I truly enjoy working with SharePoint, Power Platform, and SharePoint Framework (SPFx), and wanted to share that passion through step-by-step tutorials, guides, and training videos. My mission is to help you learn these technologies so you can utilize SharePoint, enhance productivity, and potentially build business solutions along the way.