That sounds memory bandwidth limited. Does the total t/s decode throughput improve by running multiple sessions in parallel?
(Note, that's total not per-session. Tok/s figures per session will initially tank since you're using the same total mem bandwidth to load incrementally more active params.)
(Note, that's total not per-session. Tok/s figures per session will initially tank since you're using the same total mem bandwidth to load incrementally more active params.)