What Can I Run?

Can I Run Ornith-1.0-9B Locally? 9B Dense on 16 GB vs a Used 3090

Used RTX 3090 24GB for Q4 plus KV headroom; 16 GB holds Q4 weights
Top Pick Used RTX 3090 24GB for Q4 plus KV headroom; 16 GB holds Q4 weights

Short answer: Ornith-1.0-9B is a dense ~9B coding SKU. Unsloth’s Q4_K_M GGUF is 5.7 GB; BF16 is 17.9 GB (Hub file listing, 2026-09-10). A 16 GB card and a used RTX 3090 (24 GB) both hold Q4 weights. The 3090 is the used-market default when you want KV headroom. This is not Ornith-35B and not a 100 GB MoE. LocalRig has not measured tok/s.

What SKU is Ornith-1.0-9B?

The lightweight dense member of the Ornith 1.0 family. Unsloth’s GGUF README (accessed 2026-09-10) states dense ~9B, ≈19 GB in bf16, designed for efficient single-GPU deployment. Official copy on ornith-ai/Ornith-1.0-9B / deepreinforce-ai/Ornith-1.0-9B describes 9B-Dense alongside 31B-Dense, 35B-MoE, and 397B-MoE. Unsloth’s GGUF base_model is deepreinforce-ai/Ornith-1.0-9B. License: MIT.

Architecture notes on the family card: hybrid attention (sliding window + global), long context 262,144. Vendor coding-agent scores (Terminal-Bench, SWE-Bench, etc.) are quality tables, not VRAM. One eval footnote uses Harbor with 48 GB RAM as a bench harness, not the GGUF inference floor.

Hugging Face Hub overview for ornith-ai/Ornith-1.0-9B listed ~1.5M parameters on 2026-09-10. That is inconsistent with a 17.9 GB BF16 GGUF. LocalRig treats that overview field as a shortfall, not a size.

How much memory do the Unsloth GGUFs take?

File sizes on unsloth/Ornith-1.0-9B-GGUF (Hub ls, 2026-09-10):

QuantFile size
Q3_K_S4.3 GB
Q4_K_S5.4 GB
Q4_K_M5.7 GB
UD-Q4_K_XL6.0 GB
Q5_K_M6.5 GB
Q6_K7.5 GB
Q8_09.5 GB
BF1617.9 GB

File size is not KV. A 5.7 GB Q4 still needs cache and runtime overhead. Unsloth’s README llama.cpp snippet uses -c 262144 as a max window, not a 16 GB promise. Keep context modest unless you have measured KV for this hybrid-attention SKU. Do not paste the generic dense estimator from the VRAM calculator as if it were this card.

16 GB vs used RTX 3090 — which card?

Q4 weights fit both. 24 GB is the LocalRig used-card default for context headroom; 16 GB is a weights-fit, not a 256K promise.

  • 16 GB (4060 Ti class): 5.7 GB Q4_K_M leaves ~10 GB for KV, CUDA, and OS. Short coding sessions fit. A filled 262k window is a shortfall — LocalRig has no llama-fit-params packet for this SKU.
  • Used 3090 24 GB: same Q4 file, more KV. BF16 (17.9 GB) also fits with modest context. This is the card LocalRig already documents for Qwen 3.8 27B Q4 at 4k–32k. Ornith 9B is an easier VRAM problem than 27B.

Browse used RTX 3090 24GB on eBay →

Check RTX 4060 Ti 16GB on Amazon → if you are shopping a new 16 GB card knowing Q4 9B is the job, not 70B.

8 GB cards: Unsloth does not publish a 3060 Ti tok/s table. A 5.7 GB file plus KV on 8 GB is a fit-boundary, not a LocalRig cell. No invented 8 GB tok/s.

llama.cpp vs Ollama vs vLLM

llama.cpp GGUF is what Unsloth documents. Ollama is named as a GGUF loader. vLLM snippets on the README target the dense 9B on a single GPU — not a tok/s packet. The GGUF README shows llama-server -hf deepreinforce-ai/Ornith-1.0-9B-GGUF --port 8000 -c 262144 as a vendor command. That is a how-to, not a 3090 measurement. LocalRig has no Ollama library tag plus dated tok/s for Ornith 9B as of 2026-09-10.

Related: Run llama.cpp on an RTX 3090. Ollama vs llama.cpp vs vLLM.

If you need to rent a 24 GB box to try before buying used: RunPod, Vast.ai, GPUMart. Do not rent an H100 for 9B Q4 unless you already wanted that node.

What a 9B dense coding SKU still cannot do

It does not replace 27B or 100 GB MoE on quality, and it does not unlock 262k context on 16 GB just because the Q4 file is 5.7 GB. Hybrid attention helps KV versus a fully global 9B, but LocalRig has no llama-fit-params dump for Ornith 9B. Treat the vendor 262,144 figure as a maximum on the card, then cap -c until you have a measurement. Reasoning traces (the card says assistant turns can open with a thinking block) inflate KV the same way they do on other reasoning SKUs — every think token is cache.

Compared with Qwen 3.8 27B Q4, Ornith 9B is the easier VRAM problem on the same used 3090. 27B remains the LocalRig dense daily driver when you want that parameter class. Compared with DeepSeek-V4-Flash-0731, Ornith 9B is a different memory planet (5.7 GB vs ~103 GB). Do not shop a Spark for 9B Q4 unless you already owned the Spark.

Q8_0 at 9.5 GB still fits 16 GB with short context; BF16 at 17.9 GB wants the 24 GB card or unified memory with little spare. UD-Q4_K_XL (6.0 GB) is Unsloth’s dynamic 4-bit sibling of Q4_K_M — still a 16 GB weights-fit. None of those files make an 8 GB card a documented LocalRig cell.

Power and PSU: a used 3090 is a 350 W-class part. See PSU for a multi-GPU AI rig if you are also planning a second card later for Ornith-35B. For 9B Q4, one 3090 is more card than the weights require; you are paying for KV, not for “will it load.”

Who this is NOT for

  • Shoppers who wanted Ornith-35B or 397B. Different SKUs, different files. This page is 9B dense only.
  • People treating Hub’s 1.5M parameter overview as the model size. The BF16 GGUF is 17.9 GB.
  • Anyone who needs LocalRig tok/s. No first-party row. No transplanted 3090 number from 27B.
  • Buyers who think 16 GB = full 262k context. Weights fit; filled context is unmeasured here.
  • Anyone collapsing this into Qwen 3.8 27B. 9B is easier VRAM; 27B is still the LocalRig dense daily driver on 24 GB.

Methodology

  • Fit numbers: Hub ls of unsloth/Ornith-1.0-9B-GGUF on 2026-09-10 (Q4_K_M 5.7 GB, Q8_0 9.5 GB, BF16 17.9 GB). File sizes, not LocalRig VRAM traces.
  • SKU identity: Unsloth GGUF README (dense ~9B, bf16 ≈19 GB) plus official Ornith 9B card. Hub 1.5M overview left as a discrepancy.
  • Speed: none first-party. Unsloth llama.cpp command cited as procedure only.
  • LocalRig first-party: none as of 2026-09-10.

Sources

Frequently Asked Questions

Can I run Ornith-1.0-9B on a used RTX 3090?

Yes at Q4. Unsloth's Q4_K_M GGUF is 5.7 GB. A 24 GB 3090 holds the weights with room for context. LocalRig has no first-party tok/s on this SKU.

Will a 16 GB card run Ornith 9B Q4?

The 5.7 GB Q4_K_M file fits 16 GB with headroom at short context. Longer windows (the card lists 262,144) still cost KV. 24 GB is the more comfortable used-card default.

Is this the same as Ornith-1.0-35B?

No. 9B is the lightweight dense SKU. 35B is a larger MoE in the same family. Do not reuse this table on 35B.

Why does Hugging Face list 1.5M parameters?

Hub overview for ornith-ai/Ornith-1.0-9B showed ~1.5M parameters on 2026-09-10. That conflicts with Unsloth's ~9B / 17.9 GB BF16 GGUF. Trust the GGUF file size, not that overview field.

How fast is it?

LocalRig has no first-party row. Unsloth documents llama.cpp/Ollama commands, not a 3090 tok/s table. Do not invent one.

Sources

  • unsloth/Ornith-1.0-9B-GGUF file sizes via Hub MCP, accessed 2026-09-10
  • Unsloth GGUF README: dense ~9B, ≈19 GB bf16, https://huggingface.co/unsloth/Ornith-1.0-9B-GGUF
  • ornith-ai/Ornith-1.0-9B and deepreinforce-ai/Ornith-1.0-9B official cards, accessed 2026-09-10