AI brief: your own AI server only pays off when it’s busy
Sooner or later someone says it: “If we ran our own AI, our data would stay here and we’d stop paying per use.” This week brought a real cost test for that idea, plus two updates for teams that already run their own AI tools. Here’s what changed on October 9 and 10, and what to do about it.
A real cost test: a rented GPU only pays off while it’s busy
On October 10, Microsoft’s Apps on Azure blog published a week-long test by Wuyi Weng: the open model Qwen3.8-27B served from one rented H100 GPU, a discounted Azure Spot VM at about US$2 an hour. In the busiest three hours, on October 9, it answered 1,434 requests whose tokens were worth US$24.35 at a public API list price, for US$4.65 of VM time. Over the whole week, the result flips: about 46 hours of VM time, counting setup, model loading and benchmarks, cost about US$98 plus disks and networking, against US$86 of tokens. And Spot capacity is cheap because Azure can take it back: seven evictions in about 40 hours.
Why it matters: the author’s verdict is that it pays off when it’s busy. At about US$8 of tokens per busy hour against US$2 of VM, the GPU has to work about one hour in four to break even (our arithmetic: 2 ÷ 8.12 = 0.25), roughly 41 hours a week if you rent it around the clock. Those busy hours were heavy, too: about 118,000 input tokens per request (168.7 million ÷ 1,434), whole documents rather than quick questions. The post says self-hosting makes less sense when bursty traffic leaves the GPU idle most of the day. That’s most offices: busy 9 to 5, quiet at night. And none of the test’s regions were Canadian: if your reason to self-host is keeping data here, a Canadian region’s price and availability are a separate question.
Open WebUI 0.12: one shared chat for the team, two-factor for everyone
Open WebUI, a chat interface teams host themselves in front of the models they choose, released version 0.12.0 on October 10. A shared chat can now switch from Clone only to Allow replies: colleagues keep writing in the same conversation, each message shows who sent it, and people can only edit or regenerate their own. Administrators can also require an authenticator app for every user. The project says the release includes security and access-control fixes and recommends updating production deployments.
Why it matters: a handoff, like a quote one person starts and another finishes, stays in one conversation instead of being copied around. But a self-hosted interface is not a local model: your text goes wherever the selected model runs.
n8n now blocks old code nodes, backups included
n8n 2.42.6, released October 9, refuses to create or edit workflows containing the deprecated Function, Function Item and LangChain Code nodes; the change notes say the first two still use an insecure sandbox. Saved workflows that contain them directly keep running, and you can still delete or migrate those nodes. The catches: importing a workflow that isn’t already on the instance now fails, including restoring a backup into a fresh instance, and these nodes stop working in chat tools and in sub-workflows loaded from JSON, a URL or a file. A setting, N8N_DEPRECATED_NODES_BLOCK=false, turns the check off.
Why it matters: the day you need a backup is the worst day to find out it won’t import.
Sources
Go further
- The one-page AI policy every small team should write (free template)Your team already uses AI. Six rules on one page (approved tools, a traffic light for data, human review, decisions about people, telling clients, one owner) and two prompts to draft and stress-test it.Read the article
- How to automate invoice processing with AI, step by stepA five-step workflow for a small business: one invoice inbox, AI extraction that never guesses, simple rules for GST, QST and duplicates, human approval, then your accounting software. Tested on a blurry photo with swapped digits.Read the article
What do these tasks cost you?
The calculator estimates in 2 minutes the time and money an automation can save you.