r/LocalLLM 7h ago

Question Engineers running open-source LLMs in production: what is the hardest part today?

/r/mlops/comments/1w4rm61/engineers_running_opensource_llms_in_production/
0 Upvotes

1 comment sorted by

1

u/Atretador ArchLinux Xeon E5 2673 V4 20C/40T 4x16Gb DDR4 2133 2xMI50 16Gb 7h ago
  1. Qwen 3.6 35B A3B - general backend and frontend work, debugging with real credentials

  2. Its cheap and fast

  3. nothing really

  4. long sessions with big context slows down on shitty hardware

  5. control, privacy and security

  6. if a runtime is faster I run it

  7. hardware prices

  8. my power cost is bout the same as a cheap sub like a gpt plus

  9. its on my table

  10. this doesnt seem to be a question for local models

  11. for cloud providers? I mostly run Opencode go with MiMo, price is just insane

  12. how would I have more control than: its on my table