I tested 9 LLMs on the exact same web-dev prompt for ~8 hours — RTX 3060 12GB results (Rate the best!)
- I’ve spent basically the last 8 hours testing different models on the exact same web-development prompt, and I finally finished.
- The whole point of this nine-hour test was that which local model matches the frontier-level intelligence at size and could fit easily in an RTX 3060-like consumer card.
- My setup: GPU: RTX 3060 12GB RAM: 16GB DDR4, single-channel OS: CachyOS (Arch Linux) Local models were run through my local llama.cpp setup.
Why it matters. Relevant to Model Efficiency. Matched Model Efficiency on: llama.cpp. Read it through that lens.