Skip to content

Models and Quants

MetalGlot is built around Google’s TranslateGemma family and GGUF quantizations published by mradermacher.

ModelBase repoPractical role
TranslateGemma 4Bgoogle/translategemma-4b-itSmallest family and the easiest place to start for local use.
TranslateGemma 12Bgoogle/translategemma-12b-itBalanced middle tier for stronger quality with higher memory cost.
TranslateGemma 27Bgoogle/translategemma-27b-itHighest-end family for systems that can carry a much larger model.

According to the Google model cards, the TranslateGemma family:

  • Is based on the Gemma 3 family.
  • Is designed for translation across 55 languages.
  • Supports text translation and image-to-text translation workflows.
  • Uses a 2K token input context.

The model cards also note that these models are intended specifically for translation tasks rather than broad general-purpose assistant behavior.

MetalGlot does not expose just one model file per size. It supports multiple GGUF quant levels so users can choose a model quality that fits their memory budget and speed target. The tables below use the currently published GGUF file sizes from the mradermacher repositories and summarize what each quant is for.

Quant repo: mradermacher/translategemma-4b-it-GGUF.

TypeApprox. disk spaceNotes
mmproj-Q8_00.7 GBMulti-modal supplement.
mmproj-f161.0 GBMulti-modal supplement.
Q2_K1.8 GBSmallest footprint with the largest quality tradeoff.
Q3_K_S2.0 GBSmaller footprint than the higher Q3 variants.
Q3_K_M2.2 GBLower quality.
Q3_K_L2.3 GBHeavier than Q3_K_M with some quality recovery.
IQ4_XS2.4 GBCompact 4-bit style option.
Q4_K_S2.5 GBFast, recommended.
Q4_K_M2.6 GBFast, recommended.
Q5_K_S2.9 GBHeavier mid-range quant.
Q5_K_M2.9 GBHeavier mid-range quant with slightly more headroom.
Q6_K3.3 GBVery good quality.
Q8_04.2 GBFast, best quality.
f167.9 GB16 bpw, overkill for most local use.

This is the easiest family to run on CPU-only systems and smaller GPUs.

Quant repo: mradermacher/translategemma-12b-it-GGUF.

TypeApprox. disk spaceNotes
mmproj-Q8_00.7 GBMulti-modal supplement.
mmproj-f161.0 GBMulti-modal supplement.
Q2_K4.9 GBSmallest footprint with the largest quality tradeoff.
Q3_K_S5.6 GBSmaller footprint than the higher Q3 variants.
Q3_K_M6.1 GBLower quality.
Q3_K_L6.6 GBHeavier than Q3_K_M with some quality recovery.
IQ4_XS6.7 GBCompact 4-bit style option.
Q4_K_S7.0 GBFast, recommended.
Q4_K_M7.4 GBFast, recommended.
Q5_K_S8.3 GBHeavier mid-range quant.
Q5_K_M8.5 GBHeavier mid-range quant with slightly more headroom.
Q6_K9.8 GBVery good quality.
Q8_012.6 GBFast, best quality.

This is often the practical middle ground between output quality and local hardware cost.

Quant repo: mradermacher/translategemma-27b-it-GGUF.

TypeApprox. disk spaceNotes
mmproj-Q8_00.7 GBMulti-modal supplement.
mmproj-f161.0 GBMulti-modal supplement.
Q2_K10.6 GBSmallest footprint with the largest quality tradeoff.
Q3_K_S12.3 GBSmaller footprint than the higher Q3 variants.
Q3_K_M13.5 GBLower quality.
Q3_K_L14.6 GBHeavier than Q3_K_M with some quality recovery.
IQ4_XS15.0 GBCompact 4-bit style option.
Q4_K_S15.8 GBFast, recommended.
Q4_K_M16.6 GBFast, recommended.
Q5_K_S18.9 GBHeavier mid-range quant.
Q5_K_M19.4 GBHeavier mid-range quant with slightly more headroom.
Q6_K22.3 GBVery good quality.
Q8_028.8 GBFast, best quality.

This is the high-end option for users who have enough VRAM or system memory and want the strongest quality tier available in the product.

The mmproj-Q8_0 and mmproj-f16 rows in the tables above are multi-modal supplement files. They matter when you want image translation support through the FastAPI service’s POST /translate_image route.

Typical quant families available in MetalGlot

Section titled “Typical quant families available in MetalGlot”

The exact inventory can vary by model family, but the supported tiers commonly include Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, IQ4_XS, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, and Q8_0.

MetalGlot exposes these different qualities because users run very different hardware. A lightweight laptop and a CUDA workstation should not be forced into the same default model choice.

GoalRecommended starting point
Smallest local footprint4B with Q4_K_S or Q4_K_M
Balanced local quality12B with Q4_K_M, Q5_K_M, or Q6_K
Highest quality local inference27B with Q4_K_M or above

The quant pages explicitly call out some of the common tradeoffs:

  • Quant families Q4_K_S and Q4_K_M are described as fast and recommended.
  • Quant Q6_K is called out as very good quality.
  • Quant Q8_0 is the fast, best-quality option among the listed GGUF quants.

MetalGlot’s model stack depends on work from both Google and mradermacher.

The mradermacher model pages describe these as static quants of the Google base models, with additional weighted or imatrix quant variants available in separate repositories.