Hacker News new | ask | show | jobs
by why_only_15 13 days ago
The correct comparison is not sonnet, but qwen3.5-27b on a cloud. Alibaba's pricing [0] is $0.20/m input $1.56/m output, so $0.72 for the prompt processing or $0.22 for the generation. Yours is still cheaper but the margins are less.

My guess is that this math gets less good with MoE (because you will be limited by VRAM, but clouds won't).

[0]: https://openrouter.ai/qwen/qwen3.5-27b

1 comments

I run 3.6 but yeah you are right