Xiaomi's open model picks one voice out of three talking at once.
Xiaomi released the free CocktailASR-1 speech model on Friday. You give it a short sample of the voice you want, and it ignores everyone else. On a three-voice test it got 12.29% of words wrong, against 4.11% with two voices.
Why it mattersMeeting notes and captions fall apart the moment two people talk over each other. Open weights mean this can run on your own machine, not a company's server.