Oh, GPT-3.5, recognizable in so many old screenshots by the green icon. If ever a popular model deserved the term "stochastic parrot", it was GPT-3.5. I wonder what percentage of people today still base their opinions of AI capabilities on their experiences with that model. That model was the only option for free ChatGPT users for the first year and a half of ChatGPT's existence, from November 2022 to May 2024.
GPT-3.5 was so nice because the safeguards were so easy to bypass and it wasn't lobotomized by RLHF. Newer models really don't hit the same spot and I'm not sure what I liked so much about that initial release. I look back at conversation logs I saved and I just can't replicate them nowadays.
It looks like the practical limit for that on my hardware is ~130 t/s, which is just okay. Makes reasoning mode quite painful though.
(llama.cpp can't seem to exceed 30 t/s, I had Fable make me custom inference. that's why it took me a day to get back to you)
GPT-3.5 would repeat my prompt back in its own words for me to confirm. This model immediately runs with it and generates a lot of output. I would definitely need to use a harness that allows editing / cutting off generations.
Something I've noticed about models that have been specifically uncensored or trained to generate content is that they aren't very engaging to chat with. Maybe this is just because they weren't trained to chat at all (base models) but, it's hard to find a chat model that's been trained properly.
Possibly somewhere amazing, but see also: https://x.com/pronounced_kyle/status/1768852493092680036