Qwen3.5 9B
Chat · Reasoning · Coding · Vision · Apache-2.0
Run LLMs locally with fewer guesses. Check popular local AI models against your hardware, then compare VRAM, memory, quantization, context length, and estimated tokens per second before you download.
Free to use. No software install. Hardware detection runs in your browser, and raw GPU details are not uploaded.
39 models evaluated
Search local AI models for chat, coding, reasoning, and vision. Compare compatibility, estimated memory, speed, runtime, and the evidence behind each result.
Arena ranking is shown for 20 models that can be matched exactly; unmatched models stay unranked rather than being guessed.
Arena · CC BY 4.0 · 2026-07-10Chat · Reasoning · Coding · Vision · Apache-2.0
Chat · Reasoning · Coding · Apache-2.0
Chat · Reasoning · Coding · Vision · Apache-2.0
Chat · Coding · Llama 3.1 Community
Chat · Coding · Reasoning · Apache-2.0
Coding · Chat · Apache-2.0
Reasoning · Chat · MIT
Chat · Reasoning · Coding · Apache-2.0
Chat · Reasoning · Apache-2.0
Reasoning · Chat · MIT
Reasoning · Chat · MIT
Chat · Reasoning · Vision · Gemma Terms
Coding · Chat · Apache-2.0
Chat · Reasoning · Coding · Apache-2.0
Chat · Reasoning · Vision · Gemma Terms
Chat · Coding · Reasoning · MIT
Coding · Chat · Apache-2.0
Chat · Coding · Vision · Apache-2.0
Chat · Reasoning · Apache-2.0
Chat · Reasoning · Coding · Apache-2.0
Chat · Vision · Gemma Terms
Chat · Reasoning · Apache-2.0
Chat · Coding · Reasoning · MIT
Chat · Llama 3.2 Community
Chat · Coding · Apache-2.0
Chat · Gemma Terms
Chat · Coding · MIT
Chat · Coding · Apache-2.0
Chat · Reasoning · Gemma Terms
Chat · Llama 3.2 Community
Chat · Coding · GLM-4 License
Chat · Reasoning · Llama 3.1 Community
Chat · Reasoning · Llama 3.3 Community
Reasoning · Chat · MIT + Llama
Chat · Coding · Reasoning · Apache-2.0
Chat · Coding · Reasoning · Modified MIT
Chat · Reasoning · Vision · Llama 4 Community
Chat · Coding · Reasoning · MiniMax Open License
Chat · Reasoning · Vision · Llama 4 Community
Performance is an estimate based on public specifications. Drivers, runtime versions, cooling, context length, and background load can change real results.
More than fit or fail
Use the free compatibility checker as an LLM VRAM calculator and decision guide—not just a hardware list. Every estimate keeps its inputs and limits visible.
Start with browser detection, choose a known GPU or Mac profile, or edit VRAM, unified memory, system RAM, and bandwidth yourself.
Estimate model weights, context KV cache, and runtime overhead instead of comparing a model file size with advertised VRAM alone.
Compare Q4, Q5, Q8, BF16, FP8, and NVFP4 while the checker explains memory tradeoffs and hardware-specific precision support.
See an estimated output speed range in tokens per second. It is a planning estimate, not a claim about a benchmark on your exact machine.
Open a result to inspect its memory breakdown, fit reason, runtime, source, catalog date, and evaluation engine version.
Automatic detection runs in your browser. Raw GPU adapter details are not uploaded, and you can always replace an approximate result manually.
Before you download
Running LLM locally takes more than matching parameter count to advertised VRAM. Check usable memory, model format, quantization, context, and runtime overhead before choosing a download.
Check my local LLM setupLeave room for the operating system, display, inference runtime, model context, and temporary buffers instead of treating every advertised gigabyte as available.
Match the model architecture and parameter count with a practical Q4, Q5, Q8, BF16, FP8, or NVFP4 build that your hardware and runtime support.
Compare estimated weights, KV cache, runtime overhead, fit status, and tokens per second before committing storage space and setup time.
How to run an LLM locally
Start with your PC, GPU, or Mac; adjust model precision and context; then use the fit, memory, and speed estimates to decide what to download.
Check my hardware nowThe checker tries WebGPU and WebGL locally. Browser privacy limits can make detection approximate, so you can select a profile or edit the hardware values.
Q4, Q5, and Q8 work broadly; BF16 needs more memory; FP8 and NVFP4 require specific NVIDIA architectures. Longer context increases KV-cache memory.
Use fit status, memory breakdown, estimated speed, runtime, source, catalog version, and engine version together—not a single VRAM number.
Local LLM FAQ
Practical answers about local LLM hardware requirements, VRAM, Mac unified memory, quantization, speed estimates, and browser-based detection.
Use the free local LLM compatibility checker before downloading model files. Compare fit, VRAM, memory, quantization, context, and estimated speed in one place.