Tag: bert

News
Is Nathan Chen's 4 Flip scored by Mixture-of-Experts? Part 1: Switch Transformers: sparse MoE models
Tutorials

Nathan Chen์˜ 4ํšŒ์ „ ํ”Œ๋ฆฝ์ด Mixture-of-Experts๋กœ ์ฑ„์ ๋ฉ๋‹ˆ๊นŒ? Part 1: Switch Transformers: ํฌ์†Œ MoE ๋ชจ๋ธ

Nathan Chen์˜ 4ํšŒ์ „ ์ ์ˆ˜์— ๋Œ€ํ•œ ์šฐ์Šค๊ฐฏ์†Œ๋ฆฌ๋กœ ํ‘œํ˜„๋˜๋Š” ์ด ํ† ๋ก ์€ Switch Transformer๋ฅผ ์‚ดํŽด๋ด…๋‹ˆ๋‹ค. Google์˜ Mixture-of-Experts ํ™•์žฅ์— ๋Œ€ํ•œ ๊น”๋”ํ•˜๊ณ  ๊ณผ๊ฐํ•˜๊ฒŒ ๋‹จ์ˆœํ™”๋œ ํ•ด์„์ž…๋‹ˆ๋‹ค.

2์›” 12, 2022ยท 1 min read
Is Nathan Chen's 4 Flip scored by Mixture-of-Experts? Part 2: GLaM:Efficient Scaling of LMs with MoE
Tutorials

Nathan Chen์˜ 4 Flip์ด Mixture-of-Experts๋กœ ์ฑ„์ ๋˜๋‚˜์š”? Part 2: GLaM:Efficient Scaling of LMs with MoE

MoE ๋ฏธ๋‹ˆ ์‹œ๋ฆฌ์ฆˆ์˜ Part 2๋Š” GLaM์œผ๋กœ ์ดˆ์ ์„ ๋งž์ถฅ๋‹ˆ๋‹ค. Google์˜ 1.2์กฐ-ํŒŒ๋ผ๋ฏธํ„ฐ Mixture-of-Experts ์–ธ์–ด ๋ชจ๋ธ๋กœ, ์ œ๋กœ-์ƒท, ์›-์ƒท, ๊ทธ๋ฆฌ๊ณ  ํ“จ-์ƒท ๋ฒค์น˜๋งˆํฌ์—์„œ GPT-3๊ณผ ๋™๋“ฑํ•˜๊ฑฐ๋‚˜ ๋Šฅ๊ฐ€ํ•˜๋ฉฐ ํ† ํฐ๋‹น ์•ฝ 8%์˜ ํŒŒ๋ผ๋ฏธํ„ฐ๋งŒ ํ™œ์„ฑํ™”ํ•ฉ๋‹ˆ๋‹ค

2์›” 12, 2022ยท 1 min read
BERT Paper Reviewed from a Speech Perspective
Tutorials

์Œ์„ฑ ๊ด€์ ์—์„œ ๊ฒ€ํ† ๋œ BERT ๋…ผ๋ฌธ

BERT๋Š” NLP์—์„œ ๊ฑฐ์˜ ์†Œ๊ฐœ๊ฐ€ ํ•„์š” ์—†์ง€๋งŒ, ์Œ์„ฑ ์—”์ง€๋‹ˆ์–ด์˜ ๊ด€์ ์—์„œ ์ฝ์œผ๋ฉด ํ‘œ์ค€ "๋งˆ์Šคํฌ๋œ ์–ธ์–ด ๋ชจ๋ธ๋ง์ด ๋ชจ๋“  ๊ฒƒ์„ ๋ฐ”๊ฟจ๋‹ค"๋Š” ์š”์•ฝ๊ณผ๋Š” ๋‹ค๋ฅธ ๊ตํ›ˆ๋“ค์„ ๋ฐœ๊ฒฌํ•ฉ๋‹ˆ๋‹ค.

2์›” 6, 2022ยท 1 min read
Transformer: Attention is All You Need and Listen, Attend and Spell -- from a Speech Perspective
Tutorials

Transformer: Attention is All You Need์™€ Listen, Attend and Spell -- ์Œ์„ฑ ๊ด€์ ์—์„œ

๋‘ ํŽธ์˜ ๋…ผ๋ฌธ, ํ•˜๋‚˜์˜ ์ด์•ผ๊ธฐ. "Attention Is All You Need"๋Š” ์„ธ๊ณ„์— Transformer๋ฅผ ์ œ๊ณตํ•˜๊ณ  ์ˆ˜์—ด ๋ชจ๋ธ๋ง์ด ์–ด๋–ป๊ฒŒ ์ˆ˜ํ–‰๋˜๋Š”์ง€๋ฅผ ์žฌ์„ค์ •ํ–ˆ๊ณ , "Listen, Attend and Spell" (LAS)์€ ์—”๋“œ-ํˆฌ-์—”๋“œ ์Œ์„ฑ์— ๋Œ€ํ•ด 1๋…„ ๋จผ์ € ๋™์ผํ•œ ์ผ์„ ํ–ˆ์Šตโ€ฆ

1์›” 29, 2022ยท 1 min read