Tag: nlp

News
[Olewave's Long Review] Efficient Training of Neural Transducer for Speech Recognition
Tutorials

[Olewave์˜ ๋กฑ ๋ฆฌ๋ทฐ] ์Œ์„ฑ ์ธ์‹์„ ์œ„ํ•œ ์‹ ๊ฒฝ ํŠธ๋žœ์Šค๋“€์„œ์˜ ํšจ์œจ์ ์ธ ํ›ˆ๋ จ

์‹ ๊ฒฝ ํŠธ๋žœ์Šค๋“€์„œ๋Š” ์ŠคํŠธ๋ฆฌ๋ฐ ASR์˜ ์ฃผ๋ ฅ์ด์ง€๋งŒ, ํšจ์œจ์ ์œผ๋กœ ํ›ˆ๋ จํ•˜๋Š” ๊ฒƒ์€ ์œ ๋ช…ํ•˜๊ฒŒ ๊ณ ํ†ต์Šค๋Ÿฝ์Šต๋‹ˆ๋‹ค: RNN-T ์†์‹ค์€ ์‹œํ€€์Šค ๊ธธ์ด์—์„œ 3์ œ๊ณฑ ๋ฉ”๋ชจ๋ฆฌ๋ฅผ ๊ฐ€์ง€๋ฉฐ, ์ •๋ ฌ ๊ฒฉ์ž๋Š” ๊ธด ๋ฐœํ™”์—์„œ ํญ๋ฐœํ•˜๊ณ , ๋‹จ์ผ-

7์›” 23, 2022ยท 1 min read
Boris Johnsonโ€™s Rise and Fall - an analysis of the mics
Tutorials

Boris Johnson์˜ ํฅ๋ง - ๋งˆ์ดํฌ์˜ ๋ถ„์„

์ •์น˜ ๋‰ด์Šค๋Š” ๋ณดํ†ต ์Œ์„ฑ AI ์—”์ง€๋‹ˆ์–ด๊ฐ€ ๊ธˆ์š”์ผ ์˜คํ›„์— ํ์— ์˜ฌ๋ ค๋†“๋Š” ๊ฒƒ์ด ์•„๋‹ˆ์ง€๋งŒ, Boris Johnson์˜ ์‚ฌ์ž„ ๋ฐœํ‘œ๋Š” ์ผ๋ จ์˜ ์Šค์บ”๋“ค์— ๋Œ€ํ•œ ์ •๋‹น์˜ ๋ฐ˜๋ž€์œผ๋กœ ๋งˆ์ดํฌ ๋ฑ…ํฌ ์•ž์—์„œ ์ „๋‹ฌ๋˜์—ˆ์Šต๋‹ˆ๋‹ค

7์›” 9, 2022ยท 1 min read
How Does the All-New Dictation in iOS 16 Work? Reveal Apple's Secret Sauce by a Speech Researcher!
Tutorials

iOS 16์˜ ์ƒˆ๋กœ์šด ๋ฐ›์•„์“ฐ๊ธฐ ๊ธฐ๋Šฅ์€ ์–ด๋–ป๊ฒŒ ์ž‘๋™ํ• ๊นŒ์š”? Apple์˜ ๋น„๊ฒฐ์„ ์Œ์„ฑ ์—ฐ๊ตฌ์›์ด ๋ฐํ˜€์ค๋‹ˆ๋‹ค!

Apple์˜ iOS 16 ๋ฐ›์•„์“ฐ๊ธฐ๋Š” ์™„์ „ํžˆ ๊ธฐ๊ธฐ์—์„œ ์‹คํ–‰๋˜๋ฉฐ, ์ด ๋‹จ์ผ ์„ค๊ณ„ ์„ ํƒ์€ ์ง€์—ฐ ์‹œ๊ฐ„, ๊ฐœ์ธ์ •๋ณด ๋ณดํ˜ธ, ๊ทธ๋ฆฌ๊ณ  ํœด๋Œ€ํฐ์—์„œ ์‹คํ–‰ ๊ฐ€๋Šฅํ•œ ASR ์•„ํ‚คํ…์ฒ˜์˜ ์ข…๋ฅ˜์— ๋Œ€ํ•œ ์—ฐ์‡„์  ์˜ํ–ฅ์„ ๋ฏธ์นฉ๋‹ˆ๋‹ค.

6์›” 18, 2022ยท 1 min read
[Long Review] Conformer: Convolution-augmented Transformer for Speech Recognition
Tutorials

[์žฅ๋ฌธ ๋ฆฌ๋ทฐ] Conformer: ์Œ์„ฑ ์ธ์‹์„ ์œ„ํ•œ Convolution-augmented Transformer

Conformer๋Š” end-to-end ASR์„ ์กฐ์šฉํžˆ ์žฅ์•…ํ•œ ๋ชจ๋ธ์ด๋ฉฐ, ๋ชจ๋“  Transformer ๋ธ”๋ก์— convolution ๋ชจ๋“ˆ์„ ์ถ”๊ฐ€ํ•˜๋Š” ๊ทธ ๋ฐฉ์‹์€ ๋‹จ์ˆœํ•˜์ง€๋งŒ ๋†€๋ž๊ฒŒ๋„ ์ž˜ ์ž‘๋™ํ•˜๋Š” ์•„์ด๋””์–ด ์ค‘ ํ•˜๋‚˜์ž…๋‹ˆ๋‹ค.

6์›” 11, 2022ยท 1 min read
[Long Review] Axial Attention in Multidimensional Transformers
Tutorials

[์žฅ๋ฌธ ๋ฆฌ๋ทฐ] Axial Attention in Multidimensional Transformers

์ด๋ฏธ์ง€ ๋˜๋Š” ๋น„๋””์˜ค์— ๋Œ€ํ•œ ์™„์ „ํ•œ self-attention์€ ๋น„์šฉ์ด ๊ณผ๋„ํ•ฉ๋‹ˆ๋‹ค. ์ด์ฐจ ์ฆ๊ฐ€ ์—†์ด Transformer์˜ ํ‘œํ˜„๋ ฅ์„ ์œ ์ง€ํ•˜๋ ค๋ฉด?

5์›” 28, 2022ยท 1 min read
[Short Review] Axial Attention in Multidimensional Transformers
Tutorials

[๋‹จ๋ฌธ ๋ฆฌ๋ทฐ] Axial Attention in Multidimensional Transformers

Axial attention์€ Transformer๋ฅผ ๊ณ ์ฐจ์› ๋ฐ์ดํ„ฐ์— ๋Œ€ํ•ด ๋‹ค๋ฃจ๊ธฐ ์‰ฝ๊ฒŒ ์œ ์ง€ํ•˜๋Š” ์šฐ์•„ํ•œ ์•„์ด๋””์–ด ์ค‘ ํ•˜๋‚˜์ž…๋‹ˆ๋‹ค: ํ•œ ๋ฒˆ์— ๋ชจ๋“  ํ”ฝ์…€์ด๋‚˜ ๋ชจ๋“  ์‹œ๊ฐ„-์ฃผํŒŒ์ˆ˜ ๋นˆ์— ์ฃผ๋ชฉํ•˜๋Š” ๋Œ€์‹ , ํ•œ ๋ฒˆ์— ํ•œ ์ถ•์„ ๋”ฐ๋ผ ์ฃผ๋ชฉํ•ฉ๋‹ˆ๋‹ค.

5์›” 28, 2022ยท 1 min read
[Long Review] Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis
Tutorials

[๋กฑ ๋ฆฌ๋ทฐ] Speaker Verification์—์„œ Multispeaker Text-To-Speech Synthesis๋กœ์˜ Transfer Learning

์ด ๋…ผ๋ฌธ์€ ์ œ๋กœ์ƒท ์Œ์„ฑ ํด๋กœ๋‹์„ ์‹ค์šฉ์ ์œผ๋กœ ๋งŒ๋“  ์—ฐ๊ตฌ์ž…๋‹ˆ๋‹ค: ์ˆ˜์ฒœ ๊ฐœ์˜ ๋…ธ์ด์ฆˆ๊ฐ€ ๋งŽ๊ณ  ์ „์‚ฌ๋ณธ์ด ์—†๋Š” ์Œ์„ฑ์œผ๋กœ ํŒ๋ณ„์  ๊ฒ€์ฆ ์ž‘์—…์—์„œ speaker encoder๋ฅผ ํ•™์Šตํ•œ ํ›„, ๊ทธ ์ž„๋ฒ ๋”ฉ์„ Tacotron 2๋กœ

5์›” 21, 2022ยท 1 min read
[Short Review] Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis
Tutorials

[์ˆ ๋ฆฌ๋ทฐ] Speaker Verification์—์„œ Multispeaker Text-To-Speech Synthesis๋กœ์˜ Transfer Learning

ํ˜„๋Œ€ ์ œ๋กœ์ƒท ์Œ์„ฑ ํด๋กœ๋‹ ๋ฌผ๊ฒฐ์„ ์‹œ์ž‘ํ•œ Google ๋…ผ๋ฌธ์€ ์•„๋ฆ„๋‹ต๊ฒŒ ๊น”๋”ํ•œ ์•„ํ‚คํ…์ฒ˜๋ฅผ ๊ฐ€์ง€๊ณ  ์žˆ์Šต๋‹ˆ๋‹ค: verification์— ํ•™์Šต๋œ speaker encoder, ํ…์ŠคํŠธ์™€ speaker ์ž„๋ฒ ๋”ฉ์„

5์›” 21, 2022ยท 1 min read
[Short Review] Wav2Seq: Pre-training Speech-to-Text Encoder-Decoder Models Using Pseudo Languages
Tutorials

[์ˆ ๋ฆฌ๋ทฐ] Wav2Seq: Pseudo Languages๋ฅผ ์‚ฌ์šฉํ•œ Speech-to-Text Encoder-Decoder Models์˜ Pre-training

wav2vec 2.0 ๊ฐ™์€ Self-supervised pretraining์€ ํ›Œ๋ฅญํ•œ speech encoder๋ฅผ ์ œ๊ณตํ–ˆ์Šต๋‹ˆ๋‹ค, ํ•˜์ง€๋งŒ ์ „์ฒด encoder-decoder ASR ์‹œ์Šคํ…œ์„ ์›ํ•˜๋ฉด, decoder๋Š” ์—ฌ์ „ํžˆ ์ฒ˜์Œ๋ถ€ํ„ฐ ์‹œ์ž‘ํ–ˆ์Šต๋‹ˆ๋‹ค.

5์›” 14, 2022ยท 1 min read
[Long Review] Towards Zero-Label Language Learning
Tutorials

[๊ธด ๋ฆฌ๋ทฐ] Towards Zero-Label Language Learning

์ธ๊ฐ„์ด ๋ ˆ์ด๋ธ”์„ ์ง€์ •ํ•œ ๋‹จ ํ•˜๋‚˜์˜ ์˜ˆ์ œ ์—†์ด ์ž‘์—…๋ณ„ NLP ๋ชจ๋ธ์„ ํ•™์Šต์‹œํ‚ฌ ์ˆ˜ ์žˆ๋‹ค๋ฉด ์–ด๋–จ๊นŒ์š”? "Towards Zero-Label Language Learning"์€ ์ด ์งˆ๋ฌธ์„ ๊ฐ•ํ•˜๊ฒŒ ์ œ๊ธฐํ•ฉ๋‹ˆ๋‹ค.

5์›” 7, 2022ยท 1 min read
[Long Review] Fully Sharded Data Parallel: faster AI training with fewer GPUs
Tutorials

[๊ธด ๋ฆฌ๋ทฐ] Fully Sharded Data Parallel: ๋” ์ ์€ GPU๋กœ ๋” ๋น ๋ฅธ AI ํ•™์Šต

์ ๋‹นํ•œ GPU ์˜ˆ์‚ฐ์œผ๋กœ ์ˆ˜์‹ญ์–ต ๋งค๊ฐœ๋ณ€์ˆ˜ ์Œ์„ฑ ๋˜๋Š” ์–ธ์–ด ๋ชจ๋ธ์„ ํ•™์Šตํ•˜๋Š” ๊ฒƒ์€ ํŒŒ์ดํ”„๋ผ์ธ ๋ณ‘๋ ฌํ™”, ํ…์„œ ๋ณ‘๋ ฌํ™”, ๋˜๋Š” ZeRO ์Šคํƒ€์ผ์˜ ์˜ตํ‹ฐ๋งˆ์ด์ € ์ƒค๋”ฉ ์ค‘์—์„œ ์„ ํƒํ•ด์•ผ ํ•˜๋Š” ๊ฒƒ์„ ์˜๋ฏธํ–ˆ์Šต๋‹ˆ๋‹ค.

4์›” 23, 2022ยท 1 min read
[Short Review] Fully Sharded Data Parallel: faster AI training with fewer GPUs
Tutorials

[์งง์€ ๋ฆฌ๋ทฐ] Fully Sharded Data Parallel: ๋” ์ ์€ GPU๋กœ ๋” ๋น ๋ฅธ AI ํ•™์Šต

Fully Sharded Data Parallel์€ ๋ชจ๋“  ์Œ์„ฑ ๋ฐ ์–ธ์–ด ํŒ€์ด ๊ฒฐ๊ตญ ๋ฌป๋Š” ์งˆ๋ฌธ์— ๋Œ€ํ•œ Meta์˜ ๋‹ต๋ณ€์ž…๋‹ˆ๋‹ค: ๋” ํฐ ๊ทœ๋ชจ์˜ GPU๋ฅผ ํ”„๋กœ๋น„์ €๋‹ํ•˜์ง€ ์•Š๊ณ  ํ•œ ์ž๋ฆฟ์ˆ˜ ๋” ํฐ ๋ชจ๋ธ์„ ํ•™์Šตํ•˜๋Š” ๋ฐฉ๋ฒ•์€ ๋ฌด์—‡์ผ๊นŒ์š”?

4์›” 23, 2022ยท 1 min read
[Long Review] Hurdles to Progress in Long Form Question Answering
Tutorials

[Long Review] ์žฅ๋ฌธ ์งˆ์˜์‘๋‹ต ์ง„์ „์˜ ์žฅ์• ๋ฌผ

์žฅ๋ฌธ ์งˆ์˜์‘๋‹ต์€ ๋ฆฌ๋”๋ณด๋“œ์—์„œ๋Š” ์ธ์ƒ์ ์œผ๋กœ ๋ณด์ด์ง€๋งŒ ์ž์„ธํžˆ ์‚ดํŽด๋ณด๋ฉด ๋ถˆ์•ˆํ•œ ๋ฒค์น˜๋งˆํฌ ์ค‘ ํ•˜๋‚˜์ž…๋‹ˆ๋‹ค.

4์›” 2, 2022ยท 1 min read
[Long Review] Finetuned Language Models Are Zero-Shot Learners
Tutorials

[๊ธด ๋ฆฌ๋ทฐ] ๋ฏธ์„ธ์กฐ์ •๋œ ์–ธ์–ด ๋ชจ๋ธ์€ ์ œ๋กœ์ƒท ํ•™์Šต์ž

InstructGPT๊ฐ€ ์ง€์‹œ์–ด ํŠœ๋‹์„ ์ผ์ƒ์ ์ธ ํ‘œํ˜„์œผ๋กœ ๋งŒ๋“ค๊ธฐ ์ „์—, Google์˜ FLAN ๋…ผ๋ฌธ์€ ์ž์—ฐ์–ด ์ง€์‹œ์–ด์— ๋Œ€ํ•œ ์ ๋‹นํ•œ ์–‘์˜ ๋‹ค์ค‘ ์ž‘์—… ๋ฏธ์„ธ์กฐ์ •์ด ํ‰๋ฒ”ํ•œ LM์„ ๋†€๋ž๋„๋ก c

3์›” 5, 2022ยท 1 min read
[Long Review] 'GShard': Scaling Giant Models with Conditional Computation and Automatic Sharding
Tutorials

[๊ธด ๋ฆฌ๋ทฐ] 'GShard': ์กฐ๊ฑด๋ถ€ ์—ฐ์‚ฐ ๋ฐ ์ž๋™ ์ƒค๋”ฉ์œผ๋กœ ๊ฑฐ๋Œ€ ๋ชจ๋ธ ํ™•์žฅ

Transformer๋ฅผ 100์–ต ๊ฐœ ์ด์ƒ์˜ ๋งค๊ฐœ๋ณ€์ˆ˜๋กœ ํ™•์žฅํ•˜๋Š” ๊ฒƒ์€ ๋น ๋ฅด๊ฒŒ ์ฒ ํ•™์ ์ด ๋ฉ๋‹ˆ๋‹ค: ๋ชจ๋“  ํ† ํฐ์— ๋Œ€ํ•ด ์ „์ฒด ๋„คํŠธ์›Œํฌ๋ฅผ ํ™œ์„ฑํ™”ํ•˜๊ฑฐ๋‚˜ ๊ฐ ํ† ํฐ์„ ์‹ค์ œ๋กœ ๋ด์•ผ ํ•˜๋Š” ์ „๋ฌธ๊ฐ€์—๊ฒŒ ๋ผ์šฐํŒ…ํ•ฉ๋‹ˆ๊นŒ?

2์›” 26, 2022ยท 1 min read
Triplets-like Russian Figure Skaters: Can Kullback-Leibler Divergence be Used Tell Their Difference?
Tutorials

์„ธ์Œ๋‘ฅ์ด์ฒ˜๋Ÿผ ๋‹ฎ์€ ๋Ÿฌ์‹œ์•„ ํ”ผ๊ฒจ ์Šค์ผ€์ดํ„ฐ๋“ค: Kullback-Leibler ๋ฐœ์‚ฐ์ด ๊ทธ๋“ค์˜ ์ฐจ์ด๋ฅผ ๊ตฌ๋ณ„ํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๊นŒ?

์Šคํ”ผ์ปค ์ ์‘์€ ์‹œํ€€์Šค-ํˆฌ-์‹œํ€€์Šค ๋ชจ๋ธ์„ ์ธ์ฝ”๋”์˜ ์–ด๋ ต๊ฒŒ ํš๋“ํ•œ ์Œํ–ฅ ํ‘œํ˜„์„ ์†์ƒ์‹œํ‚ค์ง€ ์•Š๊ณ  ์ ์‘์‹œํ‚ค๋ ค๊ณ  ์‹ค์ œ๋กœ ์‹œ๋„ํ•  ๋•Œ๊นŒ์ง€ ํ•ด๊ฒฐ๋œ ๊ฒƒ์ฒ˜๋Ÿผ ๋“ค๋ฆฌ๋Š” ASR ๋ฌธ์ œ ์ค‘ ํ•˜๋‚˜์ž…๋‹ˆ๋‹ค.

2์›” 19, 2022ยท 1 min read
Is Nathan Chen's 4 Flip scored by Mixture-of-Experts? Part 1: Switch Transformers: sparse MoE models
Tutorials

Nathan Chen์˜ 4ํšŒ์ „ ํ”Œ๋ฆฝ์ด Mixture-of-Experts๋กœ ์ฑ„์ ๋ฉ๋‹ˆ๊นŒ? Part 1: Switch Transformers: ํฌ์†Œ MoE ๋ชจ๋ธ

Nathan Chen์˜ 4ํšŒ์ „ ์ ์ˆ˜์— ๋Œ€ํ•œ ์šฐ์Šค๊ฐฏ์†Œ๋ฆฌ๋กœ ํ‘œํ˜„๋˜๋Š” ์ด ํ† ๋ก ์€ Switch Transformer๋ฅผ ์‚ดํŽด๋ด…๋‹ˆ๋‹ค. Google์˜ Mixture-of-Experts ํ™•์žฅ์— ๋Œ€ํ•œ ๊น”๋”ํ•˜๊ณ  ๊ณผ๊ฐํ•˜๊ฒŒ ๋‹จ์ˆœํ™”๋œ ํ•ด์„์ž…๋‹ˆ๋‹ค.

2์›” 12, 2022ยท 1 min read
Is Nathan Chen's 4 Flip scored by Mixture-of-Experts? Part 2: GLaM:Efficient Scaling of LMs with MoE
Tutorials

Nathan Chen์˜ 4 Flip์ด Mixture-of-Experts๋กœ ์ฑ„์ ๋˜๋‚˜์š”? Part 2: GLaM:Efficient Scaling of LMs with MoE

MoE ๋ฏธ๋‹ˆ ์‹œ๋ฆฌ์ฆˆ์˜ Part 2๋Š” GLaM์œผ๋กœ ์ดˆ์ ์„ ๋งž์ถฅ๋‹ˆ๋‹ค. Google์˜ 1.2์กฐ-ํŒŒ๋ผ๋ฏธํ„ฐ Mixture-of-Experts ์–ธ์–ด ๋ชจ๋ธ๋กœ, ์ œ๋กœ-์ƒท, ์›-์ƒท, ๊ทธ๋ฆฌ๊ณ  ํ“จ-์ƒท ๋ฒค์น˜๋งˆํฌ์—์„œ GPT-3๊ณผ ๋™๋“ฑํ•˜๊ฑฐ๋‚˜ ๋Šฅ๊ฐ€ํ•˜๋ฉฐ ํ† ํฐ๋‹น ์•ฝ 8%์˜ ํŒŒ๋ผ๋ฏธํ„ฐ๋งŒ ํ™œ์„ฑํ™”ํ•ฉ๋‹ˆ๋‹ค

2์›” 12, 2022ยท 1 min read
BERT Paper Reviewed from a Speech Perspective
Tutorials

์Œ์„ฑ ๊ด€์ ์—์„œ ๊ฒ€ํ† ๋œ BERT ๋…ผ๋ฌธ

BERT๋Š” NLP์—์„œ ๊ฑฐ์˜ ์†Œ๊ฐœ๊ฐ€ ํ•„์š” ์—†์ง€๋งŒ, ์Œ์„ฑ ์—”์ง€๋‹ˆ์–ด์˜ ๊ด€์ ์—์„œ ์ฝ์œผ๋ฉด ํ‘œ์ค€ "๋งˆ์Šคํฌ๋œ ์–ธ์–ด ๋ชจ๋ธ๋ง์ด ๋ชจ๋“  ๊ฒƒ์„ ๋ฐ”๊ฟจ๋‹ค"๋Š” ์š”์•ฝ๊ณผ๋Š” ๋‹ค๋ฅธ ๊ตํ›ˆ๋“ค์„ ๋ฐœ๊ฒฌํ•ฉ๋‹ˆ๋‹ค.

2์›” 6, 2022ยท 1 min read
Transformer: Attention is All You Need and Listen, Attend and Spell -- from a Speech Perspective
Tutorials

Transformer: Attention is All You Need์™€ Listen, Attend and Spell -- ์Œ์„ฑ ๊ด€์ ์—์„œ

๋‘ ํŽธ์˜ ๋…ผ๋ฌธ, ํ•˜๋‚˜์˜ ์ด์•ผ๊ธฐ. "Attention Is All You Need"๋Š” ์„ธ๊ณ„์— Transformer๋ฅผ ์ œ๊ณตํ•˜๊ณ  ์ˆ˜์—ด ๋ชจ๋ธ๋ง์ด ์–ด๋–ป๊ฒŒ ์ˆ˜ํ–‰๋˜๋Š”์ง€๋ฅผ ์žฌ์„ค์ •ํ–ˆ๊ณ , "Listen, Attend and Spell" (LAS)์€ ์—”๋“œ-ํˆฌ-์—”๋“œ ์Œ์„ฑ์— ๋Œ€ํ•ด 1๋…„ ๋จผ์ € ๋™์ผํ•œ ์ผ์„ ํ–ˆ์Šตโ€ฆ

1์›” 29, 2022ยท 1 min read
Exploring Wav2vec 2.0 fine-tuning for improved speech emotion recognition
Tutorials

์Œ์„ฑ ๊ฐ์ • ์ธ์‹ ๊ฐœ์„ ์„ ์œ„ํ•œ Wav2vec 2.0 ๋ฏธ์„ธ ์กฐ์ • ํƒ์ƒ‰

์Œ์„ฑ ๊ฐ์ • ์ธ์‹์€ ์˜ค๋žซ๋™์•ˆ ์ž‘์€ ๋ ˆ์ด๋ธ”๋œ ๋ฐ์ดํ„ฐ์…‹๊ณผ ์†์œผ๋กœ ๋งŒ๋“  ์Œํ–ฅ ํŠน์„ฑ์— ๊ฐ‡ํ˜€ ์žˆ์—ˆ์Šต๋‹ˆ๋‹ค.

12์›” 4, 2021ยท 1 min read