Skip to content
Drix10 Blog

AI Voice Language App Launch by @whosamberella

, 6 items in AI Generated Music and Audio, 3 min read

In this digest (6 items)

@whosamberella built a language app for practicing real conversations with AI voices. The app gives each character its own personality. It includes pronunciation cards for beginners.

Key points

  • Feature: Every character has its own personality.

  • Feature: Comes with pronunciation cards for beginners.

Sources

Suno launches Albums feature

Suno introduces Albums, letting users combine songs into a release with artwork, tracklist, and publishing control. Existing playlists can be converted to Albums without rebuilding.

Key points

  • Feature: Albums are live on Suno.

  • Conversion: Playlists can be turned into Albums without rebuilding.

Sources

Teleprompter Follow Voice Feature Announced

The teleprompter now includes a Follow Voice mode. It scrolls at the speaker’s pace, pauses when the speaker pauses, and waits for off‑script speech before continuing. The post also notes that scripts have been merged.

Key points

  • Follow Voice scrolls at the pace you speak.

  • Follow Voice pauses on speaker pause and resumes when speaking continues.

Sources

Songhunt multimodal AI selects gameplay‑responsive music

Songhunt introduced a multimodal AI that continuously interprets gameplay. The system selects music that matches on‑screen action. The soundtrack adapts from quiet exploration to fast‑paced combat.

Key points

  • Multimodal AI interprets gameplay continuously.

  • Soundtrack shifts with the session from quiet exploration to fast‑paced combat.

Sources

Ultrasonic Signatures Boost Instrument Recognition

The paper introduces ultrasonic signatures as a feature for musical instrument recognition. It uses 96‑kSPS recordings covering 15 instrument classes. Experiments show that ultrasonic and broadband transients improve both isolated and polyphonic recognition.

Key points

  • 96‑kSPS multitrack corpus used for evaluation.

  • Corpus includes 15 instrument classes across performers, studios, and sessions.

Sources

Indic-CLAP – Multilingual Text‑Audio Model for Nine Indic Languages

Indic-CLAP is a text‑audio model that adds a multilingual text encoder to CLAP for nine Indic languages. The audio encoder remains frozen while the new encoder is trained by distillation using machine‑translated AudioCaps captions. Evaluation shows it inherits cross‑modal alignment and performs on retrieval and zero‑shot classification tasks.

Key points

  • Model extends CLAP to nine Indic languages spoken by over a billion people.

  • Training uses distillation from CLAP’s text encoder, audio encoder, and a hybrid objective.

Sources

This digest is also a plain Markdown file in the ai-resources repository on GitHub.

All 43 in AI Generated Music and Audio