Skip to content

MarieDeVox/monologue-female-vocal-dataset

Repository files navigation

YAML

language:

  • en license: cc-by-nc-nd-4.0 tags:
  • audio
  • speech
  • voice
  • conversational
  • whisper
  • openai-whisper
  • scraped
  • web-scraped
  • in-the-wild
  • wild
  • crowdsourced
  • raw-audio
  • fine-tuning
  • alignment
  • automatic-speech-recognition
  • asr
  • speech-to-text
  • stt
  • spontaneous-speech
  • unscripted pretty_name: Conversational Female Vocal Dataset (Preview) size_categories:
  • n<10

Conversational Female Vocal Dataset (Commercial Preview)

Looking for the Full Dataset?

This Hugging Face repository contains a 3-sample preview to demonstrate acoustic quality, cadence, and data formatting. To purchase the full 32-minute, 32-file unscripted dataset with complete commercial licensing, visit the official storefront:

https://payhip.com/MarieDeVox


Overview

This dataset provides premium, high-fidelity, unscripted human speech designed for machine learning research, automated speech recognition (ASR) benchmarking, and data pipeline testing. This dataset is completely risk-free, as the vendor is the sole creator of the complete package (including voice, production, and data curation).

Unlike rigid studio scripts, this data captures authentic human cadence, spontaneous speech patterns, realistic breath placement, and natural velocity variance. The speaker evaluates real-world conversational themes surrounding relationships, self-growth, and personal development.

Available Tiers on Storefront:

  • Tier 1 (Academic/Personal): Full 32-file raw audio dataset for non-commercial research and coding practice.
  • Tier 2 (Standard Commercial): Full 32-file raw audio dataset with a commercial EULA for indie devs and software integration.
  • Tier 3 (Enterprise Bundle): Full 32-file audio dataset + Master Transcript File (.txt) + Multi-seat corporate license.

Technical Specifications

  • Format: Lossless WAV (PCM)
  • Sample Rate: Broadcast quality (44.1 kHz / 48 kHz compatible)
  • Bit Depth: 24-bit
  • Audio Preprocessing: Applied gentle high-pass filtering (80 Hz) to eliminate subsonic rumble, light noise-floor cleanup to ensure acoustic clarity, and strict peak normalization at -3.0 dB to maximize dynamic headroom.
  • Data Architecture: Pre-segmented into precise 1-minute blocks to optimize GPU Video RAM (VRAM) consumption during model ingestion and training.

About

This repository serves as the official open-source evaluation hub for a premium, high-fidelity Conversational Female Monologue Dataset. This data addresses the critical shortage of natural human velocity, spontaneous breath placement, and unscripted vocal cadence in traditional training corpora.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

No releases published

Packages

 
 
 

Contributors