You are right, the server handled connections fine. The issue was a single unindexed correlated subquery blocking the entire PostgreSQL connection pool for 13 minutes. We live and learn.
Yes, did not mean to imply that everyone should know everything immediately, just that "the server was too small" feels a bit misguided. It might solve the problem to make the hw a real monster, but if you hit limits at 500 then I think one should look at the other parts of the equation. Glad you fixed it in the end.