easy-dataset logo

easy-dataset

Free

A powerful tool for creating datasets for LLM fine-tuning 、RAG and Eval

Model APIsFreeFree tier
Inputs: text
Type
Open Source

About easy-dataset

Easy Dataset is an open-source tool for creating high-quality datasets for LLM fine-tuning, RAG, and evaluation. It addresses common challenges such as poor AI-generated QA pairs from large documents, context window limitations, duplicate questions, and the need for structured dataset management. The tool offers a full pipeline from document parsing to dataset construction, annotation, export, and evaluation. It features a project-based workflow, intelligent document chunking (chapter-aware recursive splitting), AI-assisted domain label generation, batch question and answer generation with COT support, quality check mechanisms, and multi-format export (Alpaca, ShareGPT). It also integrates a dataset plaza aggregating sources like HuggingFace and Kaggle.

Key Features

Project-based workflow for full dataset pipeline
Intelligent document processing with chapter-aware recursive chunking
AI-powered domain label system (auto-generates two-level domain trees)
Batch question generation based on text block semantics
Answer construction with support for chain-of-thought (COT) generation
Quality check mechanisms: batch delete, manual edit, AI optimization
Multi-format dataset export (Alpaca, ShareGPT) with custom field mapping
Dataset plaza aggregating HuggingFace, Kaggle, and other sources
Model configuration center supporting OpenAI API and local models (Ollama)

Pros & Cons

Pros
  • Tackles common dataset creation pain points like chunking and duplicate questions
  • Supports both API-based and local models for flexible deployment
  • Automates domain label generation to reduce manual effort
  • Enables batch generation with resume capability for large scale tasks
  • Integrates multiple data sources via dataset plaza
Cons
  • User interface and documentation currently in Chinese only
  • Requires technical setup for local model integration (Ollama)
  • May have a learning curve for non-technical users unfamiliar with dataset pipelines

Best For

Creating high-quality fine-tuning datasets for domain-specific LLMsBuilding evaluation datasets for LLM benchmarkingPreparing RAG-ready datasets from documentsConverting between dataset formats (e.g., Alpaca to ShareGPT)Generating supervised fine-tuning data with chain-of-thought reasoningManaging and annotating existing datasets with quality control

Alternatives to easy-dataset