Falcon 40B was the world’s top-ranked open-source AI model when launched. Falcon has 40 billion parameters and was trained on one trillion tokens. For two months following its launch, Falcon 40B ranked #1 on Hugging Face’s leaderboard for open source large language models (LLMs). Offered completely royalty-free with weights, Falcon 40B is revolutionary and helps democratize AI and make it a more inclusive technology.
The multilingual Falcon 40B LLM works well with English, German, Spanish, French, Italian, Portuguese, Polish, Dutch, Romanian, Czech, and Swedish languages. The foundational LLM serves as a versatile base model that can be fine-tuned for specific requirements or objectives.
Falcon 40B launched a Call for Proposals from scientists, researchers, and innovators for inspiring use cases and applications with the most exceptional use cases to receive an investment of training computing power to work on the powerful model to shape transformative solutions. The model uses only 75 percent of GPT-3’s training compute, 40 percent of Chinchilla AI’s, and 80 percent of PaLM-62B’s.
One of the core differences in the development of Falcon was the quality of the training data. The size of the pre-training data collected for Falcon 40B was nearly five trillion tokens gathered from public web crawls (~80%), research papers, legal text, news, literature, and social media conversations.
Since LLMs are particularly sensitive to the data they are trained on, our team built a custom data pipeline to extract high-quality pre-training data using extensive filtering and deduplication, implemented both at the sample level and at the string level.