A helpful place to start. One item that I would point out is that many authorities exist at the time of running the program rather than in the configuration or tool definition. So what happens on the server (e.g., what data it requests, if it sends sampling data back to the client, etc.) may not be observed by looking at the static manifest. Were you able to score these based on the declared schema, and were you able to confirm those scores by actually running the servers and observing what they request? The gap (between the declared authority and the observed authority) is where many of the risks associated with the project are likely to exist
Thanks for flagging — can you share which filters you tried? The leaderboard at capframe.ai/leaderboard filters by severity/rule client-side. If something's broken on Kubuntu/Firefox I want to fix it.