Audiobox addresses the growing need for accessible audio creation tools by combining voice inputs with text prompts. Developed by Meta's FAIR team, this research-focused platform lets users generate custom voices, sound effects, and audio stories through AI. Unlike basic text-to-speech tools, it allows voice cloning by recording samples and supports detailed sound effect creation via descriptive text. The system uses specialized models (Speech and Sound variants) built on a shared self-supervised learning foundation, offering both versatility and focused functionality. While ideal for audio creators experimenting with AI capabilities and researchers studying generative audio models, it currently emphasizes experimental use over commercial applications. Users can create and download audio stories through the Audiobox Maker demo, though format options and real-world usage limitations aren't fully detailed. Meta highlights their commitment to AI safety, making this a controlled environment to explore audio generation's possibilities.