Link: blog.neurips.cc/2024/11/27/a...
Link: blog.neurips.cc/2024/11/27/a...
"On 2017-06-12... the "Attention is all you need" paper. At the time, the focus of the research was on improving seq2seq for machine translation...."
"On 2017-06-12... the "Attention is all you need" paper. At the time, the focus of the research was on improving seq2seq for machine translation...."
I guess one could draw a line at "generating text message by message = virtuous ML" and "whole sessions = wretched GenAI" but that's just silly to me.
I guess one could draw a line at "generating text message by message = virtuous ML" and "whole sessions = wretched GenAI" but that's just silly to me.
This is insane! Hats off to him!
This is insane! Hats off to him!
- Nothing will ever be blind anymore because there's enough data in the wild to train a reviewer de-anonymizer
- What about fine-tuning a seq2seq model conditioned
- Nothing will ever be blind anymore because there's enough data in the wild to train a reviewer de-anonymizer
- What about fine-tuning a seq2seq model conditioned
From @ricardolezama.com, stay abreast of AI developments and follow.
ricardolezama.com/english/data...
From @ricardolezama.com, stay abreast of AI developments and follow.
ricardolezama.com/english/data...
i.e. “we built the modern internet”
i.e. “we built the modern internet”
LLMs are built on the Transformer architecture, a deep learning architecture based on the "Attention" algorithm
There are 3 types of transformers:
1) encoders
2) decoders
3) Seq2Seq (Encoder - Decoder)
More LLMs are typically decoder-based models
LLMs are built on the Transformer architecture, a deep learning architecture based on the "Attention" algorithm
There are 3 types of transformers:
1) encoders
2) decoders
3) Seq2Seq (Encoder - Decoder)
More LLMs are typically decoder-based models
Read about it here: doi.org/10.1162/coli... #NLProc @aixinsg.bsky.social
Read about it here: doi.org/10.1162/coli... #NLProc @aixinsg.bsky.social
The point of "Attention Is All Your Need" is that you could rip out recurrences and convolutions from seq2seq-like architectures and get better performance.
The point of "Attention Is All Your Need" is that you could rip out recurrences and convolutions from seq2seq-like architectures and get better performance.
Seq2Seq->ResNet->Transformer->S4・Hyena・Mamba
という順に詳しく解説して、実装していくのが堅そう。とか抜かす超人しかいねえ。
Seq2Seq->ResNet->Transformer->S4・Hyena・Mamba
という順に詳しく解説して、実装していくのが堅そう。とか抜かす超人しかいねえ。