<aside>

Hello 👋, I’m Lena Shakurova, AI Advisor & CEO & Founder of ParsLabs.

I’ve spent 8+ years building Conversational AI systems and helped 100+ global companies globally design, launch, and improve Conversational AI products.

Today, I help teams build reliable LLM applications, including setting up evaluation workflows so they can measure quality, catch regressions, and ship to production with confidence.

🌎 Based in Amsterdam | Connect on LinkedIn

</aside>

<aside> 🕑

Last updated on 20.07.2026

I will try to keep this resource updated as I get to know about new eval tools. New tools will be marked “WIP” until I get time to test them

📌 If you want another tool to be added send me an email to [email protected]

</aside>

<aside> 👌

If you need extra help with LLM evals, I offer audits, consultations, and full setup support to help your team build a proper LLM evaluation framework, so you can release to production with more confidence.

Send me a DM on LinkedIn or book a free intro call and I’ll send you more details about the LLM eval setup service.

</aside>

LLM evals tools: A curated library

Choosing the right evaluation tool is harder than it should be. There are dozens of frameworks for prompt testing, regression testing, LLM-as-a-judge, human evaluation, moderation, monitoring, and production observability.

This directory brings the most useful LLM evaluation tools together in one place. I've grouped them by category, noted which ones are open source, and regularly update the list as new tools appear.

Whether you're building a chatbot, AI agent, voice assistant, or another LLM application, the goal is the same: measure quality, catch regressions early, and ship with confidence.

This library includes:

  1. No-code tools for LLM evaluation (including open source)
  2. Python libraries
  3. Moderation checks
  4. Voice evaluation tools
  5. Additional resources

1. No-code tools for LLM evaluation

<aside> 💡

Check “Gallery” view for screenshots and “Open Source” to see all open source tools.

</aside>

No-code tools for LLM evaluation

<aside> 👌

If you need extra help with LLM evals, I offer:

All focused on helping your team release to production with more confidence.

📧 Send me a DM on LinkedIn or book a free intro call and I’ll send you more details our LLM eval setup service :)

</aside>

2. Python libraries for LLM evaluation

Python libraries

3. Moderation checks

Moderation checks

4. Voice evaluation tools

Voice evaluation tools

5. Additional resources

Additional resources

Does you team need help setting up LLM evaluation workflow?

<aside> 👌

If you need extra help with LLM evals, I offer:

All focused on helping your team release to production with more confidence.

📧 Send me a DM on LinkedIn or book a free intro call and I’ll send you more details our LLM eval setup service.

</aside>