Microsoft’s VASA-1 can deepfake a person with one photo and one audio track (arstechnica.com)

submitted 7 months ago by floofloof@lemmy.ca to c/artificial_intel@lemmy.ml

11 comments fedilink hide all child comments

cross-posted from: https://sh.itjust.works/post/18066953

On Tuesday, Microsoft Research Asia unveiled VASA-1, an AI model that can create a synchronized animated video of a person talking or singing from a single photo and an existing audio track. In the future, it could power virtual avatars that render locally and don't require video feeds—or allow anyone with similar tools to take a photo of a person found online and make them appear to say whatever they want.

you are viewing a single comment's thread
view the rest of the comments

[-] ech@lemm.ee 3 points 7 months ago

It's weird to always see these dismissals about how easy it is to pinpoint generated media, like we haven't already seen an insane jump in ability in just the last year. There is no future where this tech doesn't start to become a problem with its realism, and personally I think it's much closer than most seem to think it is.

this post was submitted on 19 Apr 2024

29 points (93.9% liked)

AI

4006 readers

1 users here now

Artificial intelligence (AI) is intelligence demonstrated by machines, unlike the natural intelligence displayed by humans and animals, which involves consciousness and emotionality. The distinction between the former and the latter categories is often revealed by the acronym chosen.

founded 3 years ago