Arize Phoenix
Found in: Yigtwxx/Awesome-RAG-Production
LangSmith
LangSmith is an all-in-one platform to develop, debug, test, and monitor LLM applications, ensuring high performance, accuracy, and reliability.
LangSmith is an all-in-one platform to develop, debug, test, and monitor LLM applications, ensuring high performance, accuracy, and reliability.
Langfuse
An open-source LLM engineering platform for tracing, evaluation, prompt management, and metrics. [#opensource](https://github.com/langfuse/langfuse)
An open-source LLM engineering platform for tracing, evaluation, prompt management, and metrics. #opensource Found in: steven2358/awesome-generative-ai
Helicone
Self-host / Cloud
Self-host / Cloud Found in: Yigtwxx/Awesome-RAG-Production
Weights & Biases Weave
Trace, evaluate, and improve your LLM applications
Weave by Weights & Biases is an LLM observability toolkit that provides tracing, evaluation, and dataset management for AI applications. It integrates with the broader W&B experiment tracking ecosystem.
Openllmetry
Open-source observability for your LLM application, based on OpenTelemetry 
Open-source observability for your LLM application, based on OpenTelemetry !GitHub Repo stars Found in: kyrolabs/awesome-langchain
parea.ai
Parea AI is an experimentation and human annotation platform designed for AI teams. It provides tools for experiment tracking, observability, and human annotation, helping teams confidently ship LLM applications to production. Parea AI offers features such as auto-creating domain-specific evals, performance testing and tracking, debugging failures, human review, prompt playground, deployment tools, observability, and dataset management.
Honeycomb AI
Observability platform with AI-powered debugging for LLM apps
Honeycomb extends its observability platform to LLM applications with OpenTelemetry-based tracing, natural language querying, and AI-powered root cause analysis for understanding complex AI system behavior.