| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by BikeShuester 649 days ago
	I'd suggest offering at least one free query to allow users to evaluate the service.

1 comments

rushingcreek 649 days ago

Our fast model, Phind Instant, is completely free

link

johndough 649 days ago

Maybe OP was referring to Phind-405B (the model from the article). I certainly wonder how good the 405B model really is.

link

cjtrowbridge 649 days ago

It's just an innovated (enshittified) version of Facebook's free 405b model.

link

fshr 649 days ago

Why not let us try the new model for free like the 5 uses available for the 70B model? Seems like a no brainer to hook new users if what you're selling is worth it, eh?

link

swyx 649 days ago

> The model, based on Meta Llama 3.1 8B, runs on a Phind-customized NVIDIA TensorRT-LLM inference server that offers extremely fast speeds on H100 GPUs. We start by running the model in FP8, and also enable flash decoding and fused CUDA kernels for MLP.

as far as i know you are running your own GPUs - what do you do in overload? have a queue system? what do you do in underload? just eat the costs? is there a "serverless" system here that makes sense/is anyone working on one?

link

rushingcreek 649 days ago

We run the nodes "hot" and close to overload for peak throughput. That's why NVIDIA's XQA innovation was so interesting, because it allows for much higher throughput for a given latency budget: https://github.com/NVIDIA/TensorRT-LLM/blob/main/docs/source....

Serverless would make more sense if we had a significant underutilization problem.

link