looking for gpu partners | flat-rate inference for abliterated flagships
by herpsswd - Monday August 17, 2026 at 08:42 AM
#1
been quietly building smth called untame.  

tldr: flat monthly sub (39-199€) for inference on uncensored/abliterated open-weight flagships. no payg, no token anxiety, no logs. nothing hits disk, prompts live in vram only for the gen. xmr/btc preferred, email only. openai + anthropic api, so it plugs straight into claude code / codex / opencode.  
  
why abliterated:
- no guardrails / refusal loops (especially for cyber-related tasks)
- less policy leakage + moralizing
- follows operator prompts more literally
- much better for long autonomous agent runs  
  
lineup rn: deepseek v4 flash, glm-5.2, kimi k3 (all 1M ctx). glm-5.3 + qwen3.8-max once checkpoints drop n clear our bench.  
  
napkin math works on dedicated b200/b300 nodes, but cold-start rental sucks. looking for:
- idle b200/b300 or beefy h200 clusters, rev-share / committed rates?
- would u pay flat vs openrouter payg?
- what makes “no logs” credible? audit? open source gateway?  
   
no waitlist, no link, not selling yet. pure temperature check.
Reply




 Users browsing this thread: 1 Guest(s)