Allen Institute for AI

Latest AI news, models and releases from Allen Institute for AI. ['MolmoAct2', 'OLMo 2', 'OLMo 2 7B', 'OLMo 3', 'OLMo 3 7B', 'Tülu 3']

Open Source 🇺🇸

Gigatoken: BPE Tokenizer in Rust Encodes Text at 24.53 GB/s, Up to 989x Faster Than HuggingFace Tokenizers

Gigatoken, a BPE tokenizer written in Rust by Stanford PhD student Marcel Rød, achieves text encoding speeds of 24.53 GB/s on a 144-core AMD EPYC system, outperforming HuggingFace tokenizers by 989x and OpenAI's tiktoken by 681x. The library supports 23 tokenizer families and attributes its speed to a hand-optimized pretokenizer using SWAR SIMD techniques and pretoken caching.

OpenAIOpenAI Hugging FaceHugging Face Google/DeepMindGoogle/DeepMind MetaMeta Alibaba/QwenAlibaba/Qwen DeepSeekDeepSeek Moonshot AIMoonshot AI NVIDIANVIDIA MicrosoftMicrosoft Allen Institute for AIAllen Institute for AI Answer.AIAnswer.AI MistralMistral
MarkTechPost24.07 · 01:03
Fresh news