当前位置: X-MOL 学术arXiv.cs.MM › 论文详情
Our official English website, www.x-mol.net, welcomes your feedback! (Note: you will need to create a separate account there.)
Audeo: Audio Generation for a Silent Performance Video
arXiv - CS - Multimedia Pub Date : 2020-06-23 , DOI: arxiv-2006.14348
Kun Su, Xiulong Liu, Eli Shlizerman

We present a novel system that gets as an input video frames of a musician playing the piano and generates the music for that video. Generation of music from visual cues is a challenging problem and it is not clear whether it is an attainable goal at all. Our main aim in this work is to explore the plausibility of such a transformation and to identify cues and components able to carry the association of sounds with visual events. To achieve the transformation we built a full pipeline named `\textit{Audeo}' containing three components. We first translate the video frames of the keyboard and the musician hand movements into raw mechanical musical symbolic representation Piano-Roll (Roll) for each video frame which represents the keys pressed at each time step. We then adapt the Roll to be amenable for audio synthesis by including temporal correlations. This step turns out to be critical for meaningful audio generation. As a last step, we implement Midi synthesizers to generate realistic music. \textit{Audeo} converts video to audio smoothly and clearly with only a few setup constraints. We evaluate \textit{Audeo} on `in the wild' piano performance videos and obtain that their generated music is of reasonable audio quality and can be successfully recognized with high precision by popular music identification software.

中文翻译:

Audeo:静音表演视频的音频生成

我们提出了一种新颖的系统,该系统将演奏钢琴的音乐家的视频帧作为输入,并为该视频生成音乐。从视觉线索生成音乐是一个具有挑战性的问题,目前尚不清楚它是否是一个可以实现的目标。我们在这项工作中的主要目的是探索这种转换的合理性,并确定能够将声音与视觉事件关联起来的线索和组件。为了实现转换,我们构建了一个名为“\textit{Audeo}”的完整管道,其中包含三个组件。我们首先将键盘的视频帧和音乐家的手部动作转换为原始机械音乐符号表示 Piano-Roll(Roll),用于每个视频帧,代表每个时间步按下的键。然后,我们通过包含时间相关性来调整 Roll 以适应音频合成。事实证明,此步骤对于有意义的音频生成至关重要。作为最后一步,我们实现了 Midi 合成器来生成逼真的音乐。\textit{Audeo} 只需少量设置限制即可将视频流畅清晰地转换为音频。我们在“野外”钢琴演奏视频上评估 \textit{Audeo} 并获得他们生成的音乐具有合理的音频质量,并且可以被流行的音乐识别软件以高精度成功识别。
更新日期:2020-11-10
down
wechat
bug