ORIGINAL REDDIT POST
Is anyone hosting a private chatgpt using open weights models for friends and family?
If so, what hardware and software are you using to host which models, what speeds are you getting, how many users do you have? How do you make it easily accessible to non-tech savvy people?
If so, what hardware and software are you using to host which models, what speeds are you getting, how many users do you have? How do you make it easily accessible to non-tech savvy people?
Collected discussion
Can I be part of your family?
Sorry. I need money for a B300 to run Kimi K3. Not taking more family members at this time 😄
trying to. The 2x DGX Stations have a very clear goal... but... maybe I could sell a few kidneys... B300 💙
Small startup CEO, but I used to be CTO for a public company until a few years back... then became a consultant for Fortune 500 cos.
I try not to think about the cost.
Cloudflare -> OpenWebUI -> llama.cpp (Gemma4, Qwen3.6)
Great! Someone brought up a privacy concern, are the open webui chats e2e encrypted or do you have access to others' personal chats? How fast is deepseek running on the two sparks?
Are you using this setup for yourself or are multiple people using it?
Very cool, what are you using for inference and how many people are using it?
Nice, what is the backend for openwebui? llama.cpp? What hardware are you running gemma 4 on (26b or 31?) and what speeds do you get? Is concurrent use a problem?
That's tough to hear, if I was your family I would appreciate! I thought litellm was a router only, does it also do inference?
wtf.. again wtff.. that is A LOT of compute at least like 170K ish dollars worth of hardware💀
I am building my own AI hub with gemma 4 models where my family can have basic RAG for recipes, files, etc. Also with some basic vibe building blocks to create forms and dashboards, running on a 32gb lpddr5 machine with a Ryzen 7 7735HS. Runs decent so far with smart modelrouting, MTP etc etc
My man u can run Kimi k3. U may be the chosen one
I got a fully uncensored Gemma 4 model on openwebui my mom, my little brother, and little sister use.
I'm using Hermes connected to iMessage to my whole family. Everybody has their own agent. On the hardware size... *deep breath*: - Inkling running in a GH200; - GLM 5.2 in 4x RTX Pro 6000 - DeepSeek V4 Flash in 2x GB 10 - Minimax M3 running 2x GB 10 - Qwen 3.6 27b abliterated in a Mac Studio M3 Ultra 512GB - Qwen 3.6 27b NVFP4 in a single GB 10 - Laguna S 2.1 on a single RTX Pro 6000 - Gemma 4 31b on a single RTX Pro 6000 (my wife's favorite model) - Kat Coder Dev on a Strix Halo 128gb - Step 3.7 flash on a Strix Halo 128gb I'm now waiting for my 2x DGX Stations to arrive. Exxact has been terrible with their slipping deadlines. But yeah, tok/s is amazing (the Mac Studio is a bit slow though). Total users: 4. But my wife and I and the biggest users. We use our agents a TON and I use GLM 5.2 for coding all day long. And yes, I have solar. A lot of solar. Happy to answer any questions.
My only question is… wtf do you guys do for work 😭
Great setup, but what do you get out of it? What are these agents doing for you and the wife? Do you give them tool calling and full access capabilities for your personal life, or is it really for work?
the reason why this almost always fails to gain traction is because privacy. not privacy from big tech, but privacy from you. it's one thing to talk to an llm knowing your chat history is buried somewhere in a datacenter with trillion other convos. it's a totally different thing to know your chats are directly visible by a friend/family member.
Two dgx sparks running deepseek v4 flash. Using open webui. This is for my company and serving to people who have never used ai before. Open webui is super friendly
I set.up an openwebui that points to litellm that I host. That routes to my qwen3.6 35b and glm5.2 depending on the complexity. I set up my own VPN so the fam can carry on with their context when out and about. I even built an ultra low latency duplex voice assistant the you can interrupt and it keeps context just like chatgpt. It took so much work but holy hell is it good. I've done what I could to fix all the issues that come with openwebui and make it as close to chat gpt as possible. I did an incredible job and it's seriously like 90% as good... my family doesn't touch it. GPT is so good and convenient and they love the app and being able to create images and alter them on the fly which I haven't been able to recreate yet. I just can't compete, especially at $20 a month. I'm using a 5090 for qwen and 6xpro 6000 for glm5.2. Edit. I'm actually ok with this. I didn't want to share GPU anyway.
I let a family member access openwebui through the cloudflare agent on their phone. I just have a cloudflare container that tunnels the traffic. It's all free.
Don't, you will be disappointed by their neglect and unwillingness to learn new things. I tried in so many ways over the years, but not even an network wide adblocker was accepted, since they argued they were afraid to miss good deals.
My wife and I use a harness I built with llama.cpp as the inference engine (though it can also use OpenRouter). Primarily Qwen 3.6 27B, occasionally Gemma 31B, MedGemma. It has STT and TTS, responsive web app, and tools (subprocesses) and "skills". We use it as an assistant, and I use it to plan / write code when I don't have Emacs handy, or my hands are busy -- out on a walk, cooking, childcare, running errands, etc. I would expand it to family -- I know they would appreciate the privacy, speed, tool integration -- but I worry I don't have the compute. Oddly reminds me of the 90s, when there was just one PC in the house that had to be shared. I hope someday we'll look back on hosted services the way we do at mainframes. It's really a technology that has so much potential locally.
qwen 3.6 35B A3B Q6 and gemma 12B on a strix halo, and some big open weights models on cloud