Separate music stems, dialogue and individual speakers, recover speech and align lyrics through a browser workspace, APIs or SDK.
AudioShake
Explore features, practical uses and pricing below.
AudioShake separates and analyzes audio that has already been recorded. It can extract musical parts from a finished song, divide a soundtrack into dialogue and background components, isolate speakers, recover speech, and produce timed lyrics. Its central purpose is giving users more control over a mixed recording when the original individual tracks are unavailable or inconvenient to use.
That makes it different from an AI music generator. AudioShake's separation workflows work on an existing performance. A producer may need an instrumental from an old master, a localization team may need the music-and-effects bed from a finished video, or an application may need separate voices from an overlapping conversation. The desired output determines which processing model to choose.
The service offers a browser workspace through AudioShake Studio and a developer platform with APIs and SDK access. These routes should be evaluated separately. A person processing a handful of songs needs a straightforward upload-and-review workflow; a company processing an archive needs task tracking, storage, cost accounting, and an approach to reviewing exceptional files. The underlying audio operation may be similar, but the work around it is different.
The instrument stem-separation product extracts parts such as vocals, drums, bass, guitar, piano, strings, and winds from mixed music. These outputs are often called stems, although they are estimated from the finished recording rather than recovered copies of the original studio multitracks. That distinction matters when deciding how much further processing a separated part can tolerate.
Possible uses include creating an instrumental, adjusting a musical part in a new mix, preparing a remix, and exploring a recorded arrangement. Request the parts that serve the project. If a sync pitch only needs a vocal-free version, there is little reason to request every instrument simply because those models are available. A more detailed remix may justify several outputs and a broader listening review.
Import the results into the intended editing or music-production environment and compare them with the source. Listen both to isolated parts and to the combined arrangement. An artifact that is obvious in a soloed stem may be less noticeable in a mix, while an error near an exposed vocal or sustained instrument may matter greatly. Decide suitability in the context of the actual deliverable rather than from a model's marketing description.
The dialogue, music, and effects product addresses audio from film, television, online video, interviews, and other mixed material. Its purpose is to separate speech from background components so editors can work on those elements independently. This is particularly relevant when a localization team receives a finished mix without an accompanying dialogue-free track.
AudioShake's dubbing documentation describes requesting a dialogue stem and a music-and-effects stem. The former can act as a reference for the spoken content; the latter can form a background bed beneath a new voice track. The service prepares audio components for the process. It does not itself establish that a translated script or a newly recorded performance is accurate.
For a localized clip, review the background at moments where dialogue originally overlapped important effects. Compare scene transitions, room tone, crowd sounds, and musical accents. If a removed voice leaves an obvious gap or some speech remains in the background, that section needs attention before the new language track is mixed. Keep the original mix available throughout the edit so the separated version can be assessed against the source.
Multi-speaker separation produces separate audio for individual speakers, including conversations with overlapping voices. This is a different task from diarization. Diarization identifies who spoke when; separated audio gives an editor or downstream system an actual voice track to work with. AudioShake's model documentation lists both, allowing developers to choose labels, audio outputs, or a workflow using both.
For an interview archive, isolated speakers can make selective editing and further transcription easier to organize. Still check that the right voice remains associated with the right person. Cross-talk, short interruptions, and similar-sounding voices are useful review points. The separated tracks should be judged on intelligibility and faithful representation, rather than assumed to be flawless because they have labels.
Speech recovery is another operation. AudioShake provides denoising and dereverberation models, with the distinction explained in the model reference. Denoising aims to reduce interference while retaining room character; dereverberation also reduces reflections. Choose according to the intended sound. A dry narration may suit one treatment, while an interview whose environment matters may call for a more restrained approach.
Lyric transcription and alignment turns a recorded song into text associated with the performance. Timed lyrics can support karaoke displays, subtitle-like presentation, search, and music applications. The purpose is not only to obtain words, but to connect those words with the recording at the appropriate moments.
The developer guide distinguishes line-level transcription from the alignment model used for word-level timing. The model reference also describes supplying a transcript for alignment. This is useful when an approved lyric text already exists and the project needs timings rather than another independently generated version of the words.
Review proper names, repeated choruses, ad-libs, and any unusual phrasing against the approved lyric source. If an application highlights each word as it is sung, inspect the timing in that actual display. A text file that looks correct is not enough to judge a synchronized experience. AudioShake Studio also describes a browser lyric editor with editing and realignment, giving a non-developer a route to correcting the text and its relationship to the song.
AudioShake Studio is the browser-based entry point for working with individual recordings. Its current interface presents music stems, post-production separation, lyric editing, multi-speaker processing, and speech recovery. It also accepts supported video containers, which is helpful when the audio source is part of a finished video rather than a separately exported sound file.
Start by deciding what you need to hear or download from the upload. For a song, that might be a vocal and instrumental. For a documentary clip, it might be dialogue and a background bed. For a conversation, it might be a track for each speaker. Using a named deliverable makes it easier to select a workflow and evaluate the resulting files.
The Studio landing page currently labels its music-removal-and-compliance workflow as coming soon. Related music detection and removal models are documented on the developer platform, but those API capabilities should not be treated as proof that the same workflow is already available in Studio. Check the actual browser workspace before planning a manual project around a particular tool.
The developer platform uses processing tasks with one or more requested targets. An application can submit a recording and specify the audio outputs or analysis it needs. Combining targets in one task can make asset association simpler, but each target still represents processing that needs an appropriate budget and review.
Audio processing is handled asynchronously. The task-status documentation explains checking progress and retrieving results, while the platform also documents completion notifications through webhooks. A production integration should distinguish uploaded, submitted, processing, completed, and failed items. That helps users understand whether a missing output is still being created or needs attention.
For an archive, retain a clear association between the original recording, the requested targets, and each downloaded result. A title alone is often insufficient when several versions of the same performance exist. Decide which outputs are retained, who can access them, and which files require a human listening check. Those are application-design choices around the documented API, not services automatically supplied by an endpoint.
Also define what completion means for a multi-target task. A workflow expecting both dialogue and background audio needs both assets before it can move to mixing. Preserve the requested output format in your own job record, and check that the returned file is the one the next system expects. For example, a listening preview and an archive master may require different formats. If the application allows a user to request another target later, show that as additional processing rather than silently presenting it as part of the original job. This keeps audio choices and usage accounting understandable.
AudioShake also offers an SDK for integrating separation technology into software and devices. Its product information describes supported desktop and mobile platforms and options for on-device or controlled deployment. The SDK overview is the appropriate starting point for assessing what can run in the intended environment.
A local or private deployment is relevant when the recording should remain inside a defined environment or when the application must process sound as part of a live experience. That does not mean every browser feature has an identical offline equivalent. Confirm model availability, integration requirements, supported hardware, and the commercial arrangement for the specific deployment.
Developers should use a small representative recording before designing the complete product around a processing path. Check how audio enters the system, how results are represented, and what resource constraints apply. A music practice application, a live broadcast pipeline, and an archive-processing service have different tolerance for delay and different output requirements. Choose the deployment around those requirements rather than a general claim about real-time capability.
Imagine a music publisher preparing a group of older recordings for instrumental preview use. Begin with the best available authorized masters and identify the exact recording versions. Request an instrumental or the relevant vocal and accompaniment outputs from a small representative selection. Include a dense mix, an exposed vocal passage, and a track with prominent reverberation in that selection.
Listen to each result in the intended use context. Mark problem sections and decide whether they can be repaired in the usual audio-production workflow. Keep the original master intact. If synchronized lyrics are also needed, use the approved lyric text where available and review the alignment in the target player. Instrumental suitability and lyric accuracy are separate acceptance decisions.
Only then estimate the larger batch. Record the requested targets and account for repeated processing when a project changes. For automation, use stable internal recording identifiers and keep the task results linked to them. The goal is a catalog whose outputs are traceable and usable, rather than a folder of vaguely named separated files. This example is editorial workflow guidance based on AudioShake's documented capabilities.
Studio has a free starting option and paid plans, with differences in credits, supported output formats, file duration, and tool access shown on its pricing table. Enterprise arrangements include options for team and organizational needs. Check the current plan against the required deliverable, particularly when a project needs lossless exports or longer source recordings.
The API uses credit-based processing. Its billing documentation states that credits are charged by source-audio duration and target model, with duration rounded up to the next minute. Requesting several targets therefore differs from requesting one output. Rates and length limits also differ by model; do not estimate every operation using a single generic cost per recording.
Separation is an estimate of components in a mixture. It cannot recreate an original studio session with all its routing, edits, plugins, and performance metadata. A processed output may contain residual sounds or altered texture. Low-quality source files also give the system less useful information; the lyric guide specifically warns about heavily compressed sources. Finally, obtaining separated audio does not supply permission to reuse the recording. Source and output rights remain a project consideration.
Its core workflows separate and analyze existing recordings. Instrumental extraction, vocal stems, dialogue isolation, and timed lyrics all begin with recorded audio. They should not be confused with composing a new song from a prompt.
Yes. The documented dialogue and music-and-effects targets can provide separate components for localization preparation. Translation, voice recording, final mixing, and quality review remain additional steps in the production process.
No. Speaker separation outputs audio tracks, while diarization produces speaker-related timing labels. A transcript is another artifact. Check which outputs your editing or analysis workflow actually needs before selecting the processing models.
Yes, Studio provides a browser workspace and a free starting plan. Developers can use the separate API and SDK offerings. Their access, processing limits, and billing should be checked independently rather than assumed to match a Studio subscription.