与自我关注一起在波形域内拒绝言语 (Speech Denoising in the Waveform Domain with Self-Attention) - 专知论文

会员服务 ·

0

去噪 · MoDELS · state-of-the-art · 讲稿 · HTTPS ·

2022 年 7 月 7 日

Speech Denoising in the Waveform Domain with Self-Attention

翻译：与自我关注一起在波形域内拒绝言语

Zhifeng Kong,Wei Ping,Ambrish Dantrey,Bryan Catanzaro

from arxiv, Published in ICASSP 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Listen to audio samples from CleanUNet at: https://cleanunet.github.io/

In this work, we present CleanUNet, a causal speech denoising model on the raw waveform. The proposed model is based on an encoder-decoder architecture combined with several self-attention blocks to refine its bottleneck representations, which is crucial to obtain good results. The model is optimized through a set of losses defined over both waveform and multi-resolution spectrograms. The proposed method outperforms the state-of-the-art models in terms of denoised speech quality from various objective and subjective evaluation metrics. We release our code and models at https://github.com/nvidia/cleanunet.

翻译：在这项工作中,我们提出CleanUNet,这是原始波形上的因果言分解模型,拟议的模型以编码器-解密器结构为基础,加上若干自我注意块来完善其瓶颈表示方式,这对于取得良好结果至关重要,该模型通过波形和多分辨率光谱图界定的一系列损失加以优化,拟议方法在各种客观和主观评价指标中,在非名言质量方面优于最新模型。我们在https://github.com/nvidia/cleanunet上公布了我们的代码和模型。

0

相关内容

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

MPC-1乙酰化修饰调控糖代谢影响胰腺癌生长和干性特征的机制

国家自然科学基金

0+阅读 · 2015年12月31日

CuFe2O4的形貌和尺寸可控合成及催化性能研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于无线网络小车的协同控制研究

国家自然科学基金

0+阅读 · 2013年12月31日

甲基化介导沉默MicroRNA-124调控SNAI2蛋白在结直肠癌肝转移上皮间质转化中的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

脑缺血后lncRNA调控神经元生存的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

Hybrid Spectrogram and Waveform Source Separation

Arxiv

0+阅读 · 2022年8月29日

ClusTR: Exploring Efficient Self-attention via Clustering for Vision Transformers

Arxiv

0+阅读 · 2022年8月28日

Domain Generalization in Vision: A Survey

Arxiv

17+阅读 · 2021年7月18日

Bridging the Gap Between Spectral and Spatial Domains in Graph Neural Networks

Bridging the Gap Between Spectral and Spatial Domains in Graph Neural Networks

Arxiv

15+阅读 · 2020年3月26日

Learning in the Frequency Domain

Learning in the Frequency Domain

Arxiv

11+阅读 · 2020年3月12日

VIP会员

文章信息

相关主题

state-of-the-art

相关VIP内容

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

前沿人工智能趋势报告（Frontier AI Trends Report）

【AAAI2026】善始则事半功倍：基于前缀优化的大语言模型推理强化学习

Andrej Karpathy：2025 年 LLM 年度回顾（2025 LLM Year in Review）

音退化问题：基于输入操控的鲁棒语音转换综述

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Hybrid Spectrogram and Waveform Source Separation

Arxiv

0+阅读 · 2022年8月29日

ClusTR: Exploring Efficient Self-attention via Clustering for Vision Transformers

Arxiv

0+阅读 · 2022年8月28日

Domain Generalization in Vision: A Survey

Arxiv

17+阅读 · 2021年7月18日

Bridging the Gap Between Spectral and Spatial Domains in Graph Neural Networks

Bridging the Gap Between Spectral and Spatial Domains in Graph Neural Networks

Arxiv

15+阅读 · 2020年3月26日

Learning in the Frequency Domain

Learning in the Frequency Domain

Arxiv

11+阅读 · 2020年3月12日

相关基金

MPC-1乙酰化修饰调控糖代谢影响胰腺癌生长和干性特征的机制

国家自然科学基金

0+阅读 · 2015年12月31日

CuFe2O4的形貌和尺寸可控合成及催化性能研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于无线网络小车的协同控制研究

国家自然科学基金

0+阅读 · 2013年12月31日

甲基化介导沉默MicroRNA-124调控SNAI2蛋白在结直肠癌肝转移上皮间质转化中的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

脑缺血后lncRNA调控神经元生存的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

微信扫码咨询专知VIP会员