evidentlyai/evidently

▲ 2 stars today★ 7,949⑂ 939

Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.

About evidentlyai/evidently

evidentlyai/evidently is an open-source project on GitHub, mainly written in Jupyter Notebook. Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. It currently holds 7,949 stars and 939 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the Today's Trending board, currently at rank #85 with 2 new stars today.

GitHub Repository Details

Repository evidentlyai/evidently · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

Evidently

An open-source framework to evaluate, test and monitor ML and LLM-powered systems.

https://github.com/evidentlyai/evidently/blob/HEAD/PyPi Downloads https://github.com/evidentlyai/evidently/blob/HEAD/License https://github.com/evidentlyai/evidently/blob/HEAD/PyPi Evidently

Documentation | API Reference | Discord Community | Blog | Twitter | Evidently Cloud

:bar_chart: What is Evidently?

Evidently is an open-source Python library to evaluate, test, and monitor ML and LLM systems—from experiments to production.

Evidently is very modular. You can start with one-off evaluations or host a full monitoring service.

1. Reports and Test Suites

Reports compute and summarize various data, ML and LLM quality evals.

Turn any Report into a Test Suite by adding pass/fail conditions. | Reports | |--| |Report example|

2. Monitoring Dashboard

Monitoring UI service helps visualize metrics and test results over time.

You can choose:

Evidently Cloud offers a generous free tier and extra features like dataset and user management, alerting, and no-code evals. Compare OSS vs Cloud.

| Dashboard | |--| |Dashboard example|

:woman_technologist: Install Evidently

To install from PyPI:

pip install evidently
To install Evidently using the Conda installer, run:
conda install -c conda-forge evidently

:arrow_forward: Getting started

Reports

LLM evals

This is a simple Hello World. Check the Tutorials for more: LLM evaluation.

Import the necessary components:

import pandas as pd
from evidently import Report
from evidently import Dataset, DataDefinition
from evidently.descriptors import Sentiment, TextLength, Contains
from evidently.presets import TextEvals

Create a toy dataset with questions and answers.

eval_df = pd.DataFrame([
    ["What is the capital of Japan?", "The capital of Japan is Tokyo."],
    ["Who painted the Mona Lisa?", "Leonardo da Vinci."],
    ["Can you write an essay?", "I'm sorry, but I can't assist with homework."]],
                       columns=["question", "answer"])

Create an Evidently Dataset object and add descriptors: row-level evaluators. We'll check for sentiment of each response, its length and whether it contains words indicative of denial.

eval_dataset = Dataset.from_pandas(pd.DataFrame(eval_df),
data_definition=DataDefinition(),
descriptors=[
    Sentiment("answer", alias="Sentiment"),
    TextLength("answer", alias="Length"),
    Contains("answer", items=['sorry', 'apologize'], mode="any", alias="Denials")
])

You can view the dataframe with added scores:

eval_dataset.as_dataframe()

To get a summary Report to see the distribution of scores:

report = Report([
    TextEvals()
])

my_eval = report.run(eval_dataset) my_eval

my_eval.json()

my_eval.dict()

You can also choose other evaluators, including LLM-as-a-judge and configure pass/fail conditions.

Data and ML evals

This is a simple Hello World. Check the Tutorials for more: Tabular data.

Import the Report, evaluation Preset and toy tabular dataset.

import pandas as pd
from sklearn import datasets

from evidently import Report from evidently.presets import DataDriftPreset

iris_data = datasets.load_iris(as_frame=True) iris_frame = iris_data.frame

Run the Data Drift evaluation preset that will test for shift in column distributions. Take the first 60 rows of the dataframe as "current" data and the following as reference. Get the output in Jupyter notebook:

report = Report([
    DataDriftPreset(method="psi")
],
include_tests="True")
my_eval = report.run(iris_frame.iloc[:60], iris_frame.iloc[60:])
my_eval

You can also save an HTML file. You'll need to open it from the destination folder.

my_eval.save_html("file.html")

To get the output as JSON or Python dictionary:

my_eval.json()

my_eval.dict()

You can choose other Presets, create Reports from individual Metrics and configure pass/fail conditions.

Monitoring dashboard

This launches a demo project in the locally hosted Evidently UI. Sign up for Evidently Cloud to instantly get a managed version with additional features.

if you have uv you can run Evidently UI with a single command.

uv run --with evidently evidently ui --demo-projects all

If you haven’t installed uv, create a virtual environment using the standard approach.

pip install virtualenv
virtualenv venv
source venv/bin/activate

After installing Evidently (pip install evidently), run the Evidently UI with the demo projects:

evidently ui --demo-projects all

Visit localhost:8000 to access the UI.

🚦 What can you evaluate?

Evidently has 100+ built-in evals. You can also add custom ones.

Here are examples of things you can check:

| | | |:-------------------------:|:------------------------:| | 🔡 Text descriptors | 📝 LLM outputs | | Length, sentiment, toxicity, language, special symbols, regular expression matches, etc. | Semantic similarity, retrieval relevance, summarization quality, etc. with model- and LLM-based evals. | | 🛢 Data quality | 📊 Data distribution drift | | Missing values, duplicates, min-max ranges, new categorical values, correlations, etc. | 20+ statistical tests and distance metrics to compare shifts in data distribution. | | 🎯 Classification | 📈 Regression | | Accuracy, precision, recall, ROC AUC, confusion matrix, bias, etc. | MAE, ME, RMSE, error distribution, error normality, error bias, etc. | | 🗂 Ranking (inc. RAG) | 🛒 Recommendations | | NDCG, MAP, MRR, Hit Rate, etc. | Serendipity, novelty, diversity, popularity bias, etc. |

:computer: Contributions

We welcome contributions! Read the Guide to learn more.

:books: Documentation

For more examples, refer to the complete Documentation.

Browse the API Reference for detailed API documentation.

:white_check_mark: Discord Community

If you want to chat and connect, join our Discord community!

GitHub Stars & Activity

7,949Stars
939Forks
0Open issues
Jupyter NotebookLanguage

GitHub Popularity

GitHub stars7,949
Forks939
Open issues0
Primary languageJupyter Notebook
License-
Stars gained today2
Created-
Last pushed-

Trending History

Daily boardrank #85 · ▲ 2 stars

Related AI Projects

1

patchy631 / ai-engineering-hub

Jupyter Notebook★ 38,158⑂ 6,272▲ 63 stars
→
2

datawhalechina / happy-llm

Jupyter Notebook★ 34,136⑂ 3,221▲ 20 stars
→
3

QwenLM / Qwen3-VL

Jupyter Notebook★ 20,026⑂ 1,857▲ 9 stars
→
4

chiphuyen / aie-book

Jupyter Notebook★ 17,675⑂ 2,577▲ 51 stars
→
5

NVIDIA / cosmos

Jupyter Notebook★ 11,954⑂ 896▲ 11 stars
→
6

MITDeepLearning / introtodeeplearning

Jupyter Notebook★ 8,828⑂ 4,589▲ 16 stars
→
7

ed-donner / llm_engineering

Jupyter Notebook★ 7,541⑂ 7,358▲ 13 stars
→
8

gepa-ai / gepa

Jupyter Notebook★ 6,819⑂ 556▲ 21 stars
→

More AI Rankings