Skip to content
Contact

ECAD

ECAD — 18 years solving a problem that had no ready-made solution

The problem

Every time a song plays on radio or TV, someone is entitled to be paid for it. To distribute copyright royalties fairly, ECAD, Brazil's central office for music copyright collection, needed to know, across thousands of stations broadcasting at the same time all over Brazil, which work played, on which station, when it started and when it ended.

No vendor could solve this at Brazilian scale and under Brazilian conditions. Nobody knew whether it was technically feasible. It was not a software development problem. It was a research problem.

InnoVox was born to solve it.

How it started

The project was contracted as applied research with PUC-Rio. InnoVox's two founders, then PhD candidates, took on its technical core: automatic audio identification. The company was founded to take that research into production, and the relationship is still active eighteen years later.

Discovery & Diagnosis

The request was "recognize songs", but the real problem was different. Royalty collection is not based on isolated snippets. It depends on complete performances, with precise start and end times, and on telling apart very similar versions of the same work, such as recordings with a guest artist or different arrangements. A generic snippet recognizer would not solve the problem.

Proof of Concept

The first bottleneck was search at scale. Matching each audio fragment against a database of millions of fingerprints took about 24 hours of processing for every hour of audio, which ruled out any national operation.

With purpose-built search structures and processing distributed across multiple machines, that time fell to about 3 minutes per hour of audio, a reduction of roughly 480x. This was achieved without losing tolerance to signal imperfections, and identification on clean audio reached about 99% accuracy.

Engineering & Production — four generations

  1. 1st — Radio

    Our own audio identification engine, written from scratch with no third-party dependencies, integrated with ECAD's systems. It includes segmenting the start and end of each performance and telling apart variations of the same work.

  2. 2nd — Radio and TV

    More robust search techniques and retroactive refinement of identifications, which also became the standard for radio. This is where work began on a harder problem: music on TV, under soap operas, dialogue and sound effects.

  3. 3rd — Modern architecture and live music

    Migration of the entire solution to microservices, in step with the modernization of the client's architecture. A new system to identify works performed live, even when sung by another artist, in another arrangement or in another language. Classical audio fingerprinting can't solve this problem: it required orchestrating several neural networks and adapting them to the Brazilian repertoire, since the available public models were trained mostly on English-language music.

  4. 4th — Deep neural networks

    Migration of radio and TV recognition to deep neural networks. The radio process became fully automatic, with close to 100% accuracy. On TV, accuracy reached 98% across more than 16,000 programs analyzed. The live system now orchestrates about eight neural networks, combined with the classical signal processing engine to increase precision. All of it runs under continuous optimization for the real production environment.

Results

  • ~480x faster: from ~24 h to ~3 min of processing per hour of audio.
  • ~99% accuracy on clean audio since the first generation.
  • Today: close to 100% accuracy on radio, with a fully automatic process, and 98% on TV (more than 16,000 programs analyzed).
  • More than 740,000 performances identified automatically in 2024, on radio alone.
  • Scale: about 4,000 radio stations and 400 TV channels monitored. In continuous operation since 2011 (radio), 2016 (TV) and 2024 (live music).
  • Live music identification, including covers, other performers and other languages, a problem that required original research and not just existing tools.
  • Three technology paradigms crossed without stopping operations: classical signal processing, microservices architecture and deep learning.
  • 18 years of continuous partnership, still active.

What ECAD says

The partnership with InnoVox has been fundamental for the modernization of our recorded music identification processes. With the application of artificial intelligence, we managed to automate previously manual routines, achieving a high degree of precision — about 100% accuracy on radio and 98% on television. This allowed us to scale the volume of identifications, such as the more than 740 thousand performances processed automatically in 2024 in the radio segment alone. Beyond the measurable results, we highlight InnoVox's technical capacity to meet Ecad's specificities, such as the identification of short audio clips, with just 1 second duration, and full integration with our management and audit systems.

Wilder Lopes, Manager, Music Identification Systems, ECAD

What this proves

Two capabilities InnoVox offers other clients today:

  • Sensing & Signal Intelligence: turning raw signals into reliable information, at scale and under noise.
  • Applied AI & Knowledge Engineering: orchestrating and adapting AI models to a specialized domain until they work in production, not just in a demo.

And, above both, our method: start with the problem, find out how to solve it, take the solution all the way to production, and keep evolving it as technology changes.

Do you have a problem you still haven't been able to solve?

Tell us about it.