Collections
Discover the best community collections!
Collections including paper arxiv:2311.12983
-
teknium/OpenHermes-2.5-Mistral-7B
Text Generation • 7B • Updated • 5.12k • • 905 -
teknium/OpenHermes-7B
Text Generation • Updated • 53 • 14 -
TheBloke/OpenHermes-2.5-Mistral-7B-AWQ
Text Generation • 7B • Updated • 5.32k • 22 -
GAIA: a benchmark for General AI Assistants
Paper • 2311.12983 • Published • 253
-
Attention Is All You Need
Paper • 1706.03762 • Published • 141 -
Scaling Laws for Neural Language Models
Paper • 2001.08361 • Published • 12 -
Training Compute-Optimal Large Language Models
Paper • 2203.15556 • Published • 13 -
Analogy Generation by Prompting Large Language Models: A Case Study of InstructGPT
Paper • 2210.04186 • Published
-
LNS-Madam: Low-Precision Training in Logarithmic Number System using Multiplicative Weight Update
Paper • 2106.13914 • Published • 1 -
HeurAgenix: Leveraging LLMs for Solving Complex Combinatorial Optimization Challenges
Paper • 2506.15196 • Published • 3 -
Ascend HiFloat8 Format for Deep Learning
Paper • 2409.16626 • Published • 1 -
Recipes for Pre-training LLMs with MXFP8
Paper • 2506.08027 • Published • 1
-
teknium/OpenHermes-2.5-Mistral-7B
Text Generation • 7B • Updated • 5.12k • • 905 -
teknium/OpenHermes-7B
Text Generation • Updated • 53 • 14 -
TheBloke/OpenHermes-2.5-Mistral-7B-AWQ
Text Generation • 7B • Updated • 5.32k • 22 -
GAIA: a benchmark for General AI Assistants
Paper • 2311.12983 • Published • 253
-
Attention Is All You Need
Paper • 1706.03762 • Published • 141 -
Scaling Laws for Neural Language Models
Paper • 2001.08361 • Published • 12 -
Training Compute-Optimal Large Language Models
Paper • 2203.15556 • Published • 13 -
Analogy Generation by Prompting Large Language Models: A Case Study of InstructGPT
Paper • 2210.04186 • Published
-
LNS-Madam: Low-Precision Training in Logarithmic Number System using Multiplicative Weight Update
Paper • 2106.13914 • Published • 1 -
HeurAgenix: Leveraging LLMs for Solving Complex Combinatorial Optimization Challenges
Paper • 2506.15196 • Published • 3 -
Ascend HiFloat8 Format for Deep Learning
Paper • 2409.16626 • Published • 1 -
Recipes for Pre-training LLMs with MXFP8
Paper • 2506.08027 • Published • 1