Whisper large-v3
Transcribes recordings in 99 languages. For German, there are fine-tuned versions with very low error rates. Swiss German remains a challenge: even specially adapted versions get about one word wrong out of every four.
Overview · As of August 13, 2026
Tell us what you want to do and what kind of device you have. We'll let you know which models are suitable for that.
A model that summarizes things well doesn't necessarily have to be good at programming. Select what you need on the left. On the right, you'll see all the models that can do that—whether they run on your device, are hosted by a provider, or are available as a cloud subscription.
1 What should the model be capable of?
Multiple selections are allowed.
Where it's allowed to run
By default, we display all three options side by side so you can compare them.
Select the features you want the model to have on the left. Matching models will appear here.
Frequently Searched
Tell us what kind of device you have and how you use it. We'll show you what tasks you can do with it, how smoothly they run, and what other issues you might run into.
Local AI doesn't require a dedicated graphics card. A model can run using just a processor and standard RAM. What's missing isn't the capability, but the speed.
RAM is the shelf where the model is stored. The processor is a chef who works through the recipes. A graphics card consists of hundreds of little chefs doing calculations at the same time, and the graphics memory is their shelf within easy reach.
The amount of RAM determines which model will fit. The processing power determines how much patience you'll need.
Guideline values for quantized models. In addition, there is a margin for the program, the call flow, and internal calculations.
10 out of 10
This makes the tasks achievable
That leaves 11 GB for the model, which falls into the "High-Performance Laptop" category.
Just the processor. For example, read speed. You can visibly see the response being built up; it takes a few seconds for each paragraph to appear. That's enough for occasional questions, summaries, and simple document analysis.
All tasks are already feasible. More RAM allows for larger models for the same tasks, but not more tasks, and it does not result in faster processing speeds.
You can find more details about the models listed above in the model finder.
Guidelines. If multiple requests are made simultaneously, each one occupies its own session memory. The model itself is loaded only once; the session history is duplicated.
Incidentally, the term “small language model” has no definitive definition. Microsoft cites 10 billion parameters, yet at the same time refers to a model with 14 billion as “small.” The two questions above are more useful than the term itself.
The "German" column shows how reliable our statement is; it is not a made-up grade.
| Model | Expires | Storage | German | License | Look up |
|---|---|---|---|---|---|
| Qwen 3.5 | Smartphone | 0.8B starting at 2 GB · 9B around 8 GB · 27B around 20 GB | Apache 2.0 | OllamaHugging Face | |
| Qwen 3 30B-A3B | Laptop 32 GB | about 20 GB | Good at German | Apache 2.0 | OllamaHugging Face |
| Apertus | Notebook | 8B: about 6 GB · 70B: about 45 GB | Good at German | Apache 2.0 | Hugging Face |
| Gemma 4 | Smartphone | E4B: about 4 GB · 12B: about 10 GB · 31B: about 20 GB | Good at German | Apache 2.0 | OllamaHugging Face |
| gpt-oss | Laptop 32 GB | 20B, about 14 GB · 120B, over 60 GB | Good at German | Apache 2.0 | OllamaHugging Face |
| Muse Glimmer | Desktop | about 20 GB | German (unchecked) | Apache 2.0 | Hugging Face |
| Granite 4.1 | Notebook | 3B, about 3 GB · 8B, about 7 GB · 30B, about 20 GB | Apache 2.0 | OllamaHugging Face | |
| Ministral 3 | Smartphone | 3B about 3 GB · 8B about 7 GB · 14B about 11 GB | Apache 2.0 | OllamaHugging Face | |
| Phi-4 | Notebook | 3.8B (about 3 GB) · 14B (about 11 GB) | Good at German | MIT | OllamaHugging Face |
| SmolLM3 | Smartphone | about 3 GB | Apache 2.0 | Hugging Face | |
| Teuken 7B | Notebook | about 6 GB | German: weak | Apache 2.0 only in the version with the "commercial" designation and associatedrestrictions | Hugging Face |
| EuroLLM | Notebook | 9B: about 8 GB · 22B: about 16 GB | Good at German | Apache 2.0 | Hugging Face |
| Nemotron 3 Nano | Desktop | about 20 GB | German (unchecked) | NVIDIA Open ModelLicense with Conditions | OllamaHugging Face |
| DeepSeek V4 | Server | several hundred GB | Good at German | MIT | Hugging Face |
| Mistral Large 3 and Small 4 | Server | starting at about 70 GB | Good at German | Apache 2.0 | Hugging Face |
A language model generates text. It doesn't hear, it doesn't see, and it doesn't search through a database. For those tasks, there are separate, usually very small models that you combine with a language model.
Transcribes recordings in 99 languages. For German, there are fine-tuned versions with very low error rates. Swiss German remains a challenge: even specially adapted versions get about one word wrong out of every four.
Converts text into sequences of numbers to find results that are similar in meaning rather than just exact matches. Currently the most powerful freely available model of its kind for multiple languages. It does not respond on its own.
The tried-and-true alternative, designed for over a hundred languages and longer passages. The more robust choice for multilingual documents.
Reads text from photos, documents, and forms in 32 languages. Important: It is a single model, not a combination of image recognition and a language model.
Specialized in code and trained to independently read and modify files within a project. The 30B version runs on a powerful workstation.
A European alternative to Whisper that not only transcribes spoken words but also understands their meaning. Please note regarding the text-to-speech feature: This feature may not be used for commercial purposes.
Manufacturers advertise the number of supported languages. This is a claim about the training data, not a measurement. The German Federal Printing Office conducted a measurement in collaboration with the Fraunhofer Institute: 39 models were tested on German-language administrative tasks.
MÖVE Comparison of the German Federal Printing Office, 39 models tested
Qwen3, with 4 billion parameters, ranks 8th, right behind GPT-4o. A model designed for laptops performs just as well as one from a data center. Choosing the right model is more worthwhile than upgrading.
Teuken, built specifically for the EU, scored 47.7 percent on the German knowledge test, while Meta’s Llama, which is the same size, scored 61.5 percent. Being trained for German is no substitute for good training.
The comparison evaluates German administrative language and refers to the German Basic Law when assessing values. It says nothing about Swiss High German, Helvetisms, or local administrative terminology. Since there is currently no comparable Swiss test, this remains the best available benchmark.
The Swiss model ranks third, ahead of GPT-4o. However, there is still no reliable local solution for Swiss German: when transcribing the dialect, about one in four words is incorrect. Learn more about the development of Apertus in the Swiss AI Initiative.
You can't download cloud models. Nevertheless, they serve as the most important benchmark.
Only a few models have been evaluated in German at all. To date, no German performance metrics are available for Claude Opus 5, GPT-5.6, DeepSeek V4, Qwen 3.5, Gemma 4, and Mistral Large 3, even though they have been available for quite some time. We therefore show the most recent measured results for each provider and indicate which models have since been superseded. The value for Gemini 3.1 Pro is published as a rounded figure; all others are rounded to one decimal place.
Global MMLU, German section, data collected by Artificial Analysis, figures in percent
What matters is the order of magnitude, not the ranking down to tenths of a point: There is a difference of about four points between the best cloud model and the best open model, and in some cases less than two points between the cloud models themselves.
The top-of-the-line models currently available require a server. When running on a laptop, performance for demanding tasks still lags significantly behind ChatGPT. For clearly defined tasks such as summarization, the difference is barely noticeable.
For individuals and occasional use
Twenty francs a month, no setup required, and always the most powerful model. As long as there’s no sensitive data involved, that’s the pragmatic approach.
For Sensitive Data and Independence
Nothing leaves the device, no recurring costs, and it works without an internet connection. On the other hand, it offers less performance and requires a bit of setup.
For Companies and Teams
Open models from a Swiss provider. Combines practical performance with a clear legal framework, without requiring each workstation to have its own hardware.
The third option is not very well known: Infomaniak, Swisscom, and Exoscale operate open models in Swiss data centers, some in partnership with Apertus. You pay based on usage, and the data remains in the country. If you’re looking for a ready-made chatbot rather than an API, you can find these providers under “Swiss AI Chats.” The Swiss AI Landscape lists additional providers.
You can download all of these models. However, not all of them may be used for business purposes. This is the most common stumbling block in companies.
The simple rule: If a model is licensed under Apache 2.0 or MIT, you may use it for commercial purposes, modify it, and distribute it without seeking permission. This applies to most of the models here.
For all other cases, it’s worth taking a look at the license file. Some manufacturers permit only research and personal use. Others set an upper limit on the number of users. Conditions can vary even within the same model family, so be sure to check the license for the specific model, not the manufacturer’s general license. The association offers a dedicated practical package for the company’s overall AI strategy.
It works, but not in every case. These are the four most common scenarios.
Search model, then language model
The search model finds the relevant passages in your documents, and the language model uses them to formulate the answer. Two models run simultaneously, and their memory requirements add up. This is the most common architecture of all and is already built into tools like AnythingLLM or Open WebUI.
Speech recognition, then speech model
Whisper generates the text, and the language model turns it into a transcript. The two run sequentially, not simultaneously, so the memory is sufficient for the larger of the two.
A single model
There's no need for a combination here. Models like Qwen3-VL or Gemma 4 already include image recognition. If you chain an image model and a language model together, you lose quality instead of gaining it.
Small design model, then large model
A small model predicts several words in advance, while the large one checks them all at once. The quality of the responses remains the same, but the speed increases. On standard hardware, the speed increase is a factor of one and a half to two—not the multiples often promised.
One rule applies everywhere: If two models are running at the same time, their memory requirements add up in full. If they run one after the other, the space required for the larger one is sufficient.
So you can check what's written here.
Research, compilation, and implementation were carried out with the support of AI. All information was verified against the original sources linked above. Anything that could not be verified is not included on this page.
Model versions, licenses, and measurement values change quickly. Is a model missing, or is a piece of information no longer correct? Let us know, and we'll take care of it.
We are tracking the development of open models for Switzerland and will report on any changes.