SC ScholarCircle

Running local LLMs on a laptop for research — tips?

J
just_a_student · Freshman · 5 days ago · 199 views

Got access to a lab that wants me to experiment with open-weight models on consumer hardware.

What I'm working with:

  • RTX 4060 8GB (my personal machine)
  • Lab server: 2x A5000 (48GB each)
  • Budget for API calls is tight

Looking for:

  1. Best quantization for Qwen/LLaMA 3 on 8GB
  2. vLLM vs llama.cpp for throughput
  3. Any good RAG frameworks that don't need a PhD to set up

Appreciate any pointers!

0 Replies

No replies yet — be the first to respond.

Sign in with your university account to join the discussion.