
Markus Schwarzer
CEO & Co-Founder
I was browsing the Epidemic Sound website when I noticed they had added conversational search. It was so good that, for a brief moment, I was convinced they’d found a much better solution than ours. We might as well close shop, I thought.
Then I looked into what was powering it and realized, oh… that’s us.
It wasn’t a different technology. It was Cyanite’s music understanding and search working underneath an LLM. We’d spent the past months improving our tagging models, while, at the same time, the Epidemic Sound team had been building the conversational search experience their users needed on top of them. They took the foundation we provided and extended it into a product that fit their audience.
That moment made me realize where this is all heading. As we move into the agentic era, people will build entirely new ways to work with their catalogs. But none of those workflows function unless machines can understand the music and find it. Get that layer right, and the rest becomes much easier to build.
What actually changed
Across the music industry, I keep hearing the same frustration: people spend a huge part of their day on repetitive work that feels impossible to delegate because the workflow isn’t written down anywhere. The workflow is based on years of experience that lives almost entirely in one person’s head. I’ve heard them say they’d have to clone themselves to get the work done faster.
For a long time, software could only automate workflows that were explicitly defined. The moment a process depended on someone’s accumulated experience, it became much harder to delegate.
Language models changed that. Give one enough context, and it can already mimic 80–90% of a text-based workflow. It won’t get every decision right, but it doesn’t have to. If it can handle the repetitive work, that’s already enough to save hours.
I have a friend who works in creative sync at a music publisher. Every new track that comes in ends up in one of about 30 folders she created over the years, depending on whether it fits luxury brands, automotive campaigns, FMCG, uplifting music, or whatever else makes sense to her. It’s her own taxonomy. It makes perfect sense to her because she built it, but nobody else really understands it.
At first glance, this looks like exactly the kind of workflow an LLM should be able to take over. Give it enough examples of the tracks already sitting in those folders, along with the structured music data describing them, and it starts to recognize the patterns behind her decisions. When a new song arrives, it can suggest where it belongs before she ever opens the file.
That’s the clone she’d always wished she had. But that only works because the LLM isn’t making sense of the music on its own. It’s reasoning from structured descriptions of the music and the patterns they reveal.
The gap AI can’t close on its own
Language models work from the information they’re given. The more relevant that information is, the better the result. The same applies to music. Before AI can become useful, a recording has to be translated into information a machine can reason with.
But understanding a song is only the start. An AI might know that a track is energetic, guitar-driven, and around 120 BPM. It might recognize similarities with other songs. None of that explains why someone chooses that track over another.
Think back to my friend in creative sync. Her folders aren’t simply collections of songs with similar characteristics. They’re the result of years of placing music for real clients. Every track she files away adds another example of what belongs together in her world. Over time, she’s built a way of organizing music that reflects how her business works.
AI doesn’t know any of that unless you show it. If we give it the structured music data for the tracks already sitting in those folders, it can begin to see the same patterns she sees. Instead of matching songs purely because they share musical characteristics, it learns how those characteristics relate to the way her catalog is actually used. That way it can suggest where a new song belongs before she listens to it.
This is why AI-powered workflows need more than information about the recording. They also need the knowledge a catalog has accumulated over time. That knowledge helps AI support decisions that reflect how the catalog is actually used.
That was exactly what I saw on Epidemic Sound’s website. The AI wasn’t replacing the music understanding underneath. It was using that understanding together with retrieval to make the catalog searchable through conversation.
Why the rise of AI makes structured music understanding more valuable
Music understanding gives AI a reliable way to work with a catalog. It turns every file into a consistent description of the music, giving every agent and workflow the same starting point. Search, recommendations, and automation all build on that shared understanding.
As AI becomes part of more music workflows, that foundation gets reused in more places. A conversational search system can interpret a request in natural language and connect it to the right tracks. A publisher can sort new uploads into an existing taxonomy. A creative team can search for music that fits a specific brief. Different teams ask different questions, but they’re all drawing from the same catalog.
Structured music understanding forms the foundation that these systems build on. Every time a new application appears, the catalog is already prepared for it. You don’t have to reorganize your music or describe every track again because that work has already been done. You pay the cost of understanding the music once, then every future AI workflow queries the same structured representation.
What this looks like in practice
What I saw at Epidemic Sound isn’t an isolated example. The LLM handled the text understanding while Cyanite dealt with structured music understanding behind the scenes to find, compare, and refine the right tracks.
Soundstripe has taken a similar approach. Users can search in natural language then refine their request through conversation. The language model manages that back-and-forth, asking follow-up questions, interpreting intent, and adjusting the search as the conversation evolves, while Cyanite provides the music understanding that keeps the search grounded in the catalog.
BeatStars applies the same idea in a completely different workflow. Every month, hundreds of thousands of new tracks are uploaded to its marketplace. Structured audio metadata gives every new upload a consistent description, making that catalog immediately usable for search, recommendations, and future AI applications.
These companies are solving different problems, but they have all arrived at the same conclusion. AI makes the interaction more natural. A consistent understanding of the catalog makes those interactions useful.
Looking ahead
The agentic era will bring new ways of working with music. Some will change how people search and discover music. Others will automate work that still happens manually today.
Every one of those workflows depends on a machine being able to work with the music in the catalog. The better that foundational layer, the more useful AI becomes, regardless of how the interface evolves.
That’s why AI needs both structured music understanding and retrieval. Together, they give AI a reliable way to understand, search, and reason over music, so teams can build the workflows that make sense for their business.
If you’re building for the next generation of music workflows, start with structured music understanding—the foundation they’ll all depend on.

