Overview · As of August 13, 2026

Which AI model is right for your task?

Tell us what you want to do and what kind of device you have. We'll let you know which models are suitable for that.

What can each model do?

A model that summarizes things well doesn't necessarily have to be good at programming. Select what you need on the left. On the right, you'll see all the models that can do that—whether they run on your device, are hosted by a provider, or are available as a cloud subscription.

1 What should the model be capable of?

Multiple selections are allowed.

Where it's allowed to run

By default, we display all three options side by side so you can compare them.

Select the features you want the model to have on the left. Matching models will appear here.

Frequently Searched

What can you do with your hardware?

Tell us what kind of device you have and how you use it. We'll show you what tasks you can do with it, how smoothly they run, and what other issues you might run into.

Memory and processing power are two different issues

Local AI doesn't require a dedicated graphics card. A model can run using just a processor and standard RAM. What's missing isn't the capability, but the speed.

RAM is the shelf where the model is stored. The processor is a chef who works through the recipes. A graphics card consists of hundreds of little chefs doing calculations at the same time, and the graphics memory is their shelf within easy reach.

The amount of RAM determines which model will fit. The processing power determines how much patience you'll need.

  • 7 to 8 billion4 to 6 GB
  • 14 billion8 to 10 GB
  • 32 billion18 to 22 GB
  • 70 billion40 to 50 GB

Guideline values for quantized models. In addition, there is a margin for the program, the call flow, and internal calculations.

What kind of device do you have?
What are the projections?
Just the processor and RAM. This is the most common scenario, and it's enough to get started.

10 out of 10

This makes the tasks achievable

That leaves 11 GB for the model, which falls into the "High-Performance Laptop" category.

  • Software and Operating System4 GB
  • Conversation History1 GB
  • Space for the model11 GB

Just the processor. For example, read speed. You can visibly see the response being built up; it takes a few seconds for each paragraph to appear. That's enough for occasional questions, summaries, and simple document analysis.

  • Writing TextsGemma 4 12B · 10 GB
  • SummaryGemma 4 12B · 10 GB
  • TranslateEuroLLM 9B · 8 GB
  • Search Your Own DocumentsGranite 4.1 8B · 7 GB
  • ProgrammingPhi-4 14B · 11 GB
  • Multi-step thinkingPhi-4 14B · 11 GB
  • Reading Images and ScansGemma 4 12B · 10 GB
  • Transcribe SpeechWhisper large-v3 · 3 GB
  • Search in DocumentsQwen3 Embedding 8B · 8 GB
  • Using ToolsGranite 4.1 8B · 7 GB

All tasks are already feasible. More RAM allows for larger models for the same tasks, but not more tasks, and it does not result in faster processing speeds.

You can find more details about the models listed above in the model finder.

Guidelines. If multiple requests are made simultaneously, each one occupies its own session memory. The model itself is loaded only once; the session history is duplicated.

What Else This Size Can Do

  • 0.5 billionPut it in context and fill in the gaps. Not enough for a conversation.
  • 1 billionSimple summaries and basic functions. Already pushing the limits with German.
  • 3 to 4 billionThe key takeaway for the phone call: Summarize, rephrase, and draw simple conclusions.
  • 8 billionSolid general knowledge. Hardly useful on a phone anymore, but the standard on a laptop.

On a smartphone

  • Summarize and paraphrase short texts
  • Translate easily
  • Editing Dictated Text
  • Up-to-date information, since the model does not have internet access
  • Multi-step problems involving planning and calculations
  • Long documents that require a high degree of accuracy

Incidentally, the term “small language model” has no definitive definition. Microsoft cites 10 billion parameters, yet at the same time refers to a model with 14 billion as “small.” The two questions above are more useful than the term itself.

A Comparison of All Open AI Models.

The "German" column shows how reliable our statement is; it is not a made-up grade.

ModelExpiresStorageGermanLicenseLook up
Qwen 3.5Alibaba · ChinaSmartphone0.8B starting at 2 GB · 9B around 8 GB · 27B around 20 GBGerman is usefulApache 2.0OllamaHugging Face
Qwen 3 30B-A3BAlibaba · ChinaLaptop 32 GBabout 20 GBGood at GermanApache 2.0OllamaHugging Face
ApertusETH Zurich, EPFL, and CSCS · SwitzerlandNotebook8B: about 6 GB · 70B: about 45 GBGood at GermanApache 2.0Hugging Face
Gemma 4Google · USASmartphoneE4B: about 4 GB · 12B: about 10 GB · 31B: about 20 GBGood at GermanApache 2.0OllamaHugging Face
gpt-ossOpenAI · USALaptop 32 GB20B, about 14 GB · 120B, over 60 GBGood at GermanApache 2.0OllamaHugging Face
Muse GlimmerMeta · USADesktopabout 20 GBGerman (unchecked)Apache 2.0Hugging Face
Granite 4.1IBM · USANotebook3B, about 3 GB · 8B, about 7 GB · 30B, about 20 GBGerman is usefulApache 2.0OllamaHugging Face
Ministral 3Mistral AI · FranceSmartphone3B about 3 GB · 8B about 7 GB · 14B about 11 GBGerman is usefulApache 2.0OllamaHugging Face
Phi-4Microsoft · USANotebook3.8B (about 3 GB) · 14B (about 11 GB)Good at GermanMITOllamaHugging Face
SmolLM3Hugging Face · France and the U.S.Smartphoneabout 3 GBGerman is usefulApache 2.0Hugging Face
Teuken 7BFraunhofer and OpenGPT-X · GermanyNotebookabout 6 GBGerman: weakApache 2.0 only in the version with the "commercial" designation and associatedrestrictionsHugging Face
EuroLLMEuropean Research Consortium · EUNotebook9B: about 8 GB · 22B: about 16 GBGood at GermanApache 2.0Hugging Face
Nemotron 3 NanoNVIDIA · USADesktopabout 20 GBGerman (unchecked)NVIDIA Open ModelLicense with ConditionsOllamaHugging Face
DeepSeek V4DeepSeek · ChinaServerseveral hundred GBGood at GermanMITHugging Face
Mistral Large 3 and Small 4Mistral AI · FranceServerstarting at about 70 GBGood at GermanApache 2.0Hugging Face

Models that can do something different.

A language model generates text. It doesn't hear, it doesn't see, and it doesn't search through a database. For those tasks, there are separate, usually very small models that you combine with a language model.

Language in Text

Whisper large-v3

Transcribes recordings in 99 languages. For German, there are fine-tuned versions with very low error rates. Swiss German remains a challenge: even specially adapted versions get about one word wrong out of every four.

about 3 GB · Apache 2.0

Search in Documents

Qwen3 Embedding

Converts text into sequences of numbers to find results that are similar in meaning rather than just exact matches. Currently the most powerful freely available model of its kind for multiple languages. It does not respond on its own.

0.6 GB to 8 GB · Apache 2.0

Search in Documents

BGE-M3

The tried-and-true alternative, designed for over a hundred languages and longer passages. The more robust choice for multilingual documents.

about 1 GB · MIT

Reading Images and Scans

Qwen3-VL

Reads text from photos, documents, and forms in 32 languages. Important: It is a single model, not a combination of image recognition and a language model.

Starting at about 3 GB · Apache 2.0

Programming

Qwen3 Coder

Specialized in code and trained to independently read and modify files within a project. The 30B version runs on a powerful workstation.

about 19 GB · Apache 2.0

Understanding Language

Voxtral

A European alternative to Whisper that not only transcribes spoken words but also understands their meaning. Please note regarding the text-to-speech feature: This feature may not be used for commercial purposes.

3B starting at about 3 GB · Apache 2.0, but voice output is for non-commercial use only

How well do the models speak German?

Manufacturers advertise the number of supported languages. This is a claim about the training data, not a measurement. The German Federal Printing Office conducted a measurement in collaboration with the Fraunhofer Institute: 39 models were tested on German-language administrative tasks.

Answer Questions About Germany

MÖVE Comparison of the German Federal Printing Office, 39 models tested

  1. 1Mistral Large 2.10.704
  2. 2gpt-oss 120B0.703
  3. 3Apertus 70B Switzerland0.697
  4. 4Qwen3 30B-A3B0.694
  5. 5Gemma 2 9B0.693
  6. 6GPT-4:not available0.692
  7. 7Llama 3.3 70B0.691
  8. 8Qwen3 4B0.689
  9. 9DeepSeek R10.689
  10. 10Phi-40.689

Size matters less than expected

Qwen3, with 4 billion parameters, ranks 8th, right behind GPT-4o. A model designed for laptops performs just as well as one from a data center. Choosing the right model is more worthwhile than upgrading.

European origin is no guarantee

Teuken, built specifically for the EU, scored 47.7 percent on the German knowledge test, while Meta’s Llama, which is the same size, scored 61.5 percent. Being trained for German is no substitute for good training.

The measurements were taken in Germany

The comparison evaluates German administrative language and refers to the German Basic Law when assessing values. It says nothing about Swiss High German, Helvetisms, or local administrative terminology. Since there is currently no comparable Swiss test, this remains the best available benchmark.

Apertus is stronger than his reputation

The Swiss model ranks third, ahead of GPT-4o. However, there is still no reliable local solution for Swiss German: when transcribing the dialect, about one in four words is incorrect. Learn more about the development of Apertus in the Swiss AI Initiative.

How far is that from ChatGPT?

You can't download cloud models. Nevertheless, they serve as the most important benchmark.

Only a few models have been evaluated in German at all. To date, no German performance metrics are available for Claude Opus 5, GPT-5.6, DeepSeek V4, Qwen 3.5, Gemma 4, and Mistral Large 3, even though they have been available for quite some time. We therefore show the most recent measured results for each provider and indicate which models have since been superseded. The value for Gemini 3.1 Pro is published as a rounded figure; all others are rounded to one decimal place.

Global MMLU, German section, data collected by Artificial Analysis, figures in percent

Gemini 3.1Pro - Current≈95.0Cloud only
Gemini 3 Proreplaced by 3.1 Pro93.2Cloud only
Claude Opus 4.5replaced by Opus 593.0Cloud only
GPT-5.2replaced by GPT-5.691.5Cloud only
DeepSeek V3.2Replaced by V490.8open, but server
Llama 4 Maverick87.8open, but server
gpt-oss120B (current)87.3open, but server

What matters is the order of magnitude, not the ranking down to tenths of a point: There is a difference of about four points between the best cloud model and the best open model, and in some cases less than two points between the cloud models themselves.

The top-of-the-line models currently available require a server. When running on a laptop, performance for demanding tasks still lags significantly behind ChatGPT. For clearly defined tasks such as summarization, the difference is barely noticeable.

  • GPT-5.6OpenAI, USAFree tier · Plus $20/month · Pro $100/month
  • Claude Opus 5Anthropic, USAFree plan · Pro: $20/month · Max: starting at $100/month
  • Gemini 3.1 ProGoogle, USAFree tier · AI Plus CHF 5 · AI Pro CHF 17 · AI Ultra starting at CHF 100
  • MistralMistral AI, FranceFree plan · Pro: about 15 EUR/month

Cloud Subscription

For individuals and occasional use

Twenty francs a month, no setup required, and always the most powerful model. As long as there’s no sensitive data involved, that’s the pragmatic approach.

Locally on your own device

For Sensitive Data and Independence

Nothing leaves the device, no recurring costs, and it works without an internet connection. On the other hand, it offers less performance and requires a bit of setup.

Hosted in Switzerland

For Companies and Teams

Open models from a Swiss provider. Combines practical performance with a clear legal framework, without requiring each workstation to have its own hardware.

The third option is not very well known: Infomaniak, Swisscom, and Exoscale operate open models in Swiss data centers, some in partnership with Apertus. You pay based on usage, and the data remains in the country. If you’re looking for a ready-made chatbot rather than an API, you can find these providers under “Swiss AI Chats.” The Swiss AI Landscape lists additional providers.

Are you allowed to use it in the store?

You can download all of these models. However, not all of them may be used for business purposes. This is the most common stumbling block in companies.

The simple rule: If a model is licensed under Apache 2.0 or MIT, you may use it for commercial purposes, modify it, and distribute it without seeking permission. This applies to most of the models here.

For all other cases, it’s worth taking a look at the license file. Some manufacturers permit only research and personal use. Others set an upper limit on the number of users. Conditions can vary even within the same model family, so be sure to check the license for the specific model, not the manufacturer’s general license. The association offers a dedicated practical package for the company’s overall AI strategy.

These four have a catch

  • Codestral 22B is licensed under Mistral's non-production license. It is not permitted for commercial use, even though it is available for free download and is included in Ollama.
  • Command A (111B) is approved only for research and personal use. Its successor, Command A+, has been truly free since May 2026.
  • Teuken 7B, Basic Version: Not for commercial use. Only the version labeled "commercial" is licensed under Apache 2.0.
  • MiniMax M2.7 Commercial use is permitted only with the manufacturer's express permission.

Combine multiple models?

It works, but not in every case. These are the four most common scenarios.

Search Documents

Suitable for everyday use

Search model, then language model

The search model finds the relevant passages in your documents, and the language model uses them to formulate the answer. Two models run simultaneously, and their memory requirements add up. This is the most common architecture of all and is already built into tools like AnythingLLM or Open WebUI.

Take notes on the meeting and summarize it

Suitable for everyday use

Speech recognition, then speech model

Whisper generates the text, and the language model turns it into a transcript. The two run sequentially, not simultaneously, so the memory is sufficient for the larger of the two.

Review Documents and Forms

A Common Misconception

A single model

There's no need for a combination here. Models like Qwen3-VL or Gemma 4 already include image recognition. If you chain an image model and a language model together, you lose quality instead of gaining it.

Speed up responses

It works, but the effect is moderate

Small design model, then large model

A small model predicts several words in advance, while the large one checks them all at once. The quality of the responses remains the same, but the speed increases. On standard hardware, the speed increase is a factor of one and a half to two—not the multiples often promised.

One rule applies everywhere: If two models are running at the same time, their memory requirements add up in full. If they run one after the other, the space required for the larger one is sufficient.

Where the information comes from.

So you can check what's written here.

About the Creation of This Page

Research, compilation, and implementation were carried out with the support of AI. All information was verified against the original sources linked above. Anything that could not be verified is not included on this page.

Model versions, licenses, and measurement values change quickly. Is a model missing, or is a piece of information no longer correct? Let us know, and we'll take care of it.

Stay up to date.

We are tracking the development of open models for Switzerland and will report on any changes.