Models and Quants
MetalGlot is built around Google’s TranslateGemma family and GGUF quantizations published by mradermacher.
Base models
Section titled “Base models”| Model | Base repo | Practical role |
|---|---|---|
| TranslateGemma 4B | google/translategemma-4b-it | Smallest family and the easiest place to start for local use. |
| TranslateGemma 12B | google/translategemma-12b-it | Balanced middle tier for stronger quality with higher memory cost. |
| TranslateGemma 27B | google/translategemma-27b-it | Highest-end family for systems that can carry a much larger model. |
According to the Google model cards, the TranslateGemma family:
- Is based on the Gemma 3 family.
- Is designed for translation across 55 languages.
- Supports text translation and image-to-text translation workflows.
- Uses a 2K token input context.
The model cards also note that these models are intended specifically for translation tasks rather than broad general-purpose assistant behavior.
GGUF quantized variants
Section titled “GGUF quantized variants”MetalGlot does not expose just one model file per size. It supports multiple GGUF quant levels so users can choose a model quality that fits their memory budget and speed target. The tables below use the currently published GGUF file sizes from the mradermacher repositories and summarize what each quant is for.
TranslateGemma 4B GGUF
Section titled “TranslateGemma 4B GGUF”Quant repo: mradermacher/translategemma-4b-it-GGUF.
| Type | Approx. disk space | Notes |
|---|---|---|
mmproj-Q8_0 | 0.7 GB | Multi-modal supplement. |
mmproj-f16 | 1.0 GB | Multi-modal supplement. |
Q2_K | 1.8 GB | Smallest footprint with the largest quality tradeoff. |
Q3_K_S | 2.0 GB | Smaller footprint than the higher Q3 variants. |
Q3_K_M | 2.2 GB | Lower quality. |
Q3_K_L | 2.3 GB | Heavier than Q3_K_M with some quality recovery. |
IQ4_XS | 2.4 GB | Compact 4-bit style option. |
Q4_K_S | 2.5 GB | Fast, recommended. |
Q4_K_M | 2.6 GB | Fast, recommended. |
Q5_K_S | 2.9 GB | Heavier mid-range quant. |
Q5_K_M | 2.9 GB | Heavier mid-range quant with slightly more headroom. |
Q6_K | 3.3 GB | Very good quality. |
Q8_0 | 4.2 GB | Fast, best quality. |
f16 | 7.9 GB | 16 bpw, overkill for most local use. |
This is the easiest family to run on CPU-only systems and smaller GPUs.
TranslateGemma 12B GGUF
Section titled “TranslateGemma 12B GGUF”Quant repo: mradermacher/translategemma-12b-it-GGUF.
| Type | Approx. disk space | Notes |
|---|---|---|
mmproj-Q8_0 | 0.7 GB | Multi-modal supplement. |
mmproj-f16 | 1.0 GB | Multi-modal supplement. |
Q2_K | 4.9 GB | Smallest footprint with the largest quality tradeoff. |
Q3_K_S | 5.6 GB | Smaller footprint than the higher Q3 variants. |
Q3_K_M | 6.1 GB | Lower quality. |
Q3_K_L | 6.6 GB | Heavier than Q3_K_M with some quality recovery. |
IQ4_XS | 6.7 GB | Compact 4-bit style option. |
Q4_K_S | 7.0 GB | Fast, recommended. |
Q4_K_M | 7.4 GB | Fast, recommended. |
Q5_K_S | 8.3 GB | Heavier mid-range quant. |
Q5_K_M | 8.5 GB | Heavier mid-range quant with slightly more headroom. |
Q6_K | 9.8 GB | Very good quality. |
Q8_0 | 12.6 GB | Fast, best quality. |
This is often the practical middle ground between output quality and local hardware cost.
TranslateGemma 27B GGUF
Section titled “TranslateGemma 27B GGUF”Quant repo: mradermacher/translategemma-27b-it-GGUF.
| Type | Approx. disk space | Notes |
|---|---|---|
mmproj-Q8_0 | 0.7 GB | Multi-modal supplement. |
mmproj-f16 | 1.0 GB | Multi-modal supplement. |
Q2_K | 10.6 GB | Smallest footprint with the largest quality tradeoff. |
Q3_K_S | 12.3 GB | Smaller footprint than the higher Q3 variants. |
Q3_K_M | 13.5 GB | Lower quality. |
Q3_K_L | 14.6 GB | Heavier than Q3_K_M with some quality recovery. |
IQ4_XS | 15.0 GB | Compact 4-bit style option. |
Q4_K_S | 15.8 GB | Fast, recommended. |
Q4_K_M | 16.6 GB | Fast, recommended. |
Q5_K_S | 18.9 GB | Heavier mid-range quant. |
Q5_K_M | 19.4 GB | Heavier mid-range quant with slightly more headroom. |
Q6_K | 22.3 GB | Very good quality. |
Q8_0 | 28.8 GB | Fast, best quality. |
This is the high-end option for users who have enough VRAM or system memory and want the strongest quality tier available in the product.
Vision support files
Section titled “Vision support files”The mmproj-Q8_0 and mmproj-f16 rows in the tables above are multi-modal supplement files. They matter when you want image translation support through the FastAPI service’s POST /translate_image route.
Typical quant families available in MetalGlot
Section titled “Typical quant families available in MetalGlot”The exact inventory can vary by model family, but the supported tiers commonly include Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, IQ4_XS, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, and Q8_0.
MetalGlot exposes these different qualities because users run very different hardware. A lightweight laptop and a CUDA workstation should not be forced into the same default model choice.
Practical guidance
Section titled “Practical guidance”| Goal | Recommended starting point |
|---|---|
| Smallest local footprint | 4B with Q4_K_S or Q4_K_M |
| Balanced local quality | 12B with Q4_K_M, Q5_K_M, or Q6_K |
| Highest quality local inference | 27B with Q4_K_M or above |
The quant pages explicitly call out some of the common tradeoffs:
- Quant families
Q4_K_SandQ4_K_Mare described as fast and recommended. - Quant
Q6_Kis called out as very good quality. - Quant
Q8_0is the fast, best-quality option among the listed GGUF quants.
Credits and provenance
Section titled “Credits and provenance”MetalGlot’s model stack depends on work from both Google and mradermacher.
- Google provides the original TranslateGemma base models: 4B, 12B, and 27B.
- The mradermacher account provides the GGUF quantizations used for practical local deployment: 4B GGUF, 12B GGUF, and 27B GGUF.
The mradermacher model pages describe these as static quants of the Google base models, with additional weighted or imatrix quant variants available in separate repositories.