langextract
A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.
A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.
Cheetu AI
Your Lightweight Interpreter and AI Notetaker
Cheetu AI provides real-time transcription, live translation, and instant AI summaries for every meeting, lecture, or interview.
Voice Inbox
Voice Inbox is a tool designed for quickly capturing thoughts on the go. It transcribes spoken words with human-level accuracy and saves them to a journal, allowing users to focus on expressing themselves and managing tasks. It integrates with Obsidian for seamless note-taking.
ClearCypherAI
ClearCypher LLC is a company that builds Generative AI products, including Audio to Audio (T2T) speech engine, Text to Audio (T2A) speech engine, and Audio to Text (A2T) transcription engine. They offer machine learning solutions specializing in automatic speech recognition, machine translation, optical character recognition, and speaker identification. Their platform provides language technology solutions for processing audio, video, image, and text content, delivering enterprise-grade language translation and voice biometrics.
HanLP
中文分词 词性标注 命名实体识别 依存句法分析 成分句法分析 语义依存分析 语义角色标注 指代消解 风格转换 语义相似度 新词发现 关键词短语提取 自动摘要 文本分类聚类 拼音简繁转换 自然语言处理
中文分词 词性标注 命名实体识别 依存句法分析 成分句法分析 语义依存分析 语义角色标注 指代消解 风格转换 语义相似度 新词发现 关键词短语提取 自动摘要 文本分类聚类 拼音简繁转换 自然语言处理
datasets
🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
500-AI-Machine-learning-Deep-learning-Computer-vision-NLP-Projects-with-code
500 AI Machine learning Deep learning Computer vision NLP Projects with code
500 AI Machine learning Deep learning Computer vision NLP Projects with code
BettaFish
微舆:人人可用的多Agent舆情分析助手,打破信息茧房,还原舆情原貌,预测未来走向,辅助决策!从0实现,不依赖任何框架。
微舆:人人可用的多Agent舆情分析助手,打破信息茧房,还原舆情原貌,预测未来走向,辅助决策!从0实现,不依赖任何框架。
unilm
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
ailearning
AiLearning:数据分析+机器学习实战+线性代数+PyTorch+NLTK+TF2
AiLearning:数据分析+机器学习实战+线性代数+PyTorch+NLTK+TF2
Audiosonic
Transform Text into Realistic Audio with Audiosonic by Writesonic
Audiosonic is an AI-powered text-to-speech tool developed by Writesonic that transforms written text into realistic, human-like audio 1. Its core purpose is to provide high-quality, engaging audio content quickly and easily, eliminating the need for expensive voice actors and recording studios 12. Key features include: Realistic, human-like audio generation using advanced deep learning algorithms 12 Support for over 30 languages and dialects 18 Customizable voice settings (gender, accent, tone, speed, pitch) 112 Instant AI voice generation 12 Commercial use clearance for generated audio 12 Potential applications include marketing and advertising, education, podcast production, accessibility solutions, software demos, and content repurposing. Audiosonic's unique selling points are its high-quality natural-sounding audio, extensive multilingual support, ease of use, instant audio generation, and seamless integration with Writesonic 112. Technically, Audiosonic is a cloud-based SaaS application requiring an internet connection 15. It integrates fully within the Writesonic platform, streamlining the content creation process 112. While specific awards or recognition are not documented, Audiosonic was released in September 2023 and continues to be improved 612. Its advanced capabilities and integration with Writesonic position it as a powerful tool for businesses and content creators seeking to efficiently produce high-quality audio content.
CV
✅(已完结)超级全面的 深度学习 笔记【土堆 Pytorch】【李沐 动手学深度学习】【吴恩达 深度学习】【大飞 大模型Agent】
✅(已完结)超级全面的 深度学习 笔记【土堆 Pytorch】【李沐 动手学深度学习】【吴恩达 深度学习】【大飞 大模型Agent】
Awesome-Chinese-LLM
整理开源的中文大语言模型,以规模较小、可私有化部署、训练成本较低的模型为主,包括底座模型,垂直领域微调及应用,数据集与教程等。
整理开源的中文大语言模型,以规模较小、可私有化部署、训练成本较低的模型为主,包括底座模型,垂直领域微调及应用,数据集与教程等。
OpenWispr
Open source voice-to-text assistant, 3x faster than typing.
OpenWispr is an open-source, AI-powered voice dictation tool that converts your voice into formatted text instantly. It runs 100% locally, ensuring full privacy, and is designed to be 3-5x faster than typing. It's especially useful for prompting LLMs, writing emails, sending texts, and works seamlessly across various applications, allowing users to pick their preferred model and even edit the system prompt for full control.
SpeechLab
SpeechLab's Natural-Sounding AI Voice Solutions
SpeechLab provides a cutting-edge AI platform that overcomes language barriers using sophisticated speech-to-speech translation and dubbing tools. Supported by Andrew Ng’s AI Fund and leading investors, it delivers top-tier features like superior transcription, context-aware translation, and dubbed audio that sounds almost identical to human voices. Users can translate, transcribe, and dub material across various languages and dialects, achieving a flexible, detailed conveyance of ideas and feelings with lifelike accuracy. The service emphasizes ethical standards, mandating that users possess rights to any voices utilized, and follows rigorous protocols to prevent unauthorized voice cloning. Perfect for media, business, and education fields, SpeechLab fits effortlessly into current processes, offering a scalable, team-oriented platform customized for content producers, companies, and schools. Pricing options range from a free initial trial to full-service white-glove support, rendering premium dubbing and translation available to everyone.
ImbaTTS - Free unlimited Text to Speech
ImbaTTS is a free, unlimited text-to-speech tool that runs entirely in your browser, with all processing and rendering done locally. Powered by the open-source Piper TTS project, it offers natural-sounding voice synthesis in over 50 languages.
Podbrews
AI-Powered Document-to-Podcast Conversion with Podbrews
Podbrews is an innovative platform designed to convert your written documents into engaging podcast-style audio files using advanced AI technology. Users can seamlessly transform any PDF into a compelling audio experience, perfect for on-the-go listening or accessibility needs. Podbrews stands out by offering a plethora of audio styles to choose from, including sci-fi, fantasy, and public radio, ensuring that the final product matches the user's desired aesthetic. This tool is perfect for professionals, educators, and content creators looking to reach a wider audience through audio formats.
best-of-ml-python
🏆 A ranked list of awesome machine learning Python libraries. Updated weekly.
🏆 A ranked list of awesome machine learning Python libraries. Updated weekly.
RAG_Techniques
This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.
This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.
SpeechFlow - Advanced Speech-to-Text API
SpeechFlow is a multilingual Speech-to-Text API that offers state-of-the-art accuracy in 14 languages. It converts sound to text, speech to text, and audio to text with high accuracy. SpeechFlow supports both cloud and on-prem deployment.
luvvoice
Luvvoice: Free AI Text‑to‑Speech with 200+ Voices, 70+ Languages, and Voice Cloning
Luvvoice is a free online AI text‑to‑speech (TTS) platform that converts text and documents into natural‑sounding audio using real AI voices. With 200+ AI voices across 70+ languages and dialects, it supports advanced voice cloning, easy text‑to‑audio, and document‑to‑voice (including PDF). Luvvoice offers generous usage with no ads or CAPTCHA, extended character limits (up to 20,000 per conversion and 20,000,000 per month for standard voices), and flexible Free, Basic, and Pro plans—positioning it as a leading ElevenLabs alternative for 2025.
Whisper GitHub
Whisper is a general-purpose speech recognition model developed by OpenAI. It is trained on a large dataset of diverse audio and is also a multi-task model that can perform multilingual speech recognition as well as speech translation and language identification. Whisper uses a Transformer sequence-to-sequence model trained on various speech processing tasks, including multilingual speech recognition, speech translation, spoken language identification, and voice activity detection. These tasks are jointly represented as a sequence of tokens to be predicted by the decoder, allowing a single model to replace many stages of a traditional speech-processing pipeline. The multitask training format uses a set of special tokens that serve as task specifiers or classification targets.
spaCy
💫 Industrial-strength Natural Language Processing (NLP) in Python
💫 Industrial-strength Natural Language Processing (NLP) in Python
WhisperUI - Text to Speech
WhisperUI is a text to speech and speech to text service powered by OpenAI Whisper API. With WhisperUI you can use your OpenAI api keys to get affordable text to speech and speech to text services. It allows users to convert audio files to text and SRT files using OpenAI Whisper Speech to Text.