WangChanLION

Overview

WangchanLion is an instruction-finetuned model based on SEA-LION, a pan-ASEAN pretrained LLM development led by AI Singapore. The finetuning of WangchanLion is a collaborative effort between VISTEC and AI Singapore.

Solution

What WangchanLion offers:

  • Transparent pretrained model: The development of SEA-LION is community-driven, with different ASEAN collaborators contributing pretraining datasets. The SEA-LION developers ensure that all datasets are safe and can be utilized without commercial restrictions. This transparency extends to the provision of pretraining code, ensuring anyone can replicate SEA-LION using the provided datasets.
  • Transparent finetuning data: In the spirit of open science, we make the finetuning data for WangchanLion accessible to all. This commitment to openness empowers the community by providing complete visibility into the instruction finetuning data that shapes WangchanLion.
  • Transparent finetuning code: The finetuning code for WangchanLion is readily available for distribution. By sharing our methods and processes, we invite others to learn from, build upon, and innovate alongside us.

WanhchanLion is intended to be used as an instruction-following model for reading comprehension, brainstorming, and creative writing. The model is suitable to be finetuned further to support any typical instruction-following use cases.

WangchanLion consist of a range of model families, including:

  • WangchanLionv3 – WangchanLION-v3 8B model is continuously pre-trained on Llama-SEA-LION-v3-8B-IT using 47 billion high-quality Thai tokens.
  • WangchanLionv2 – Gemma2 9B WangchanLIONv2 Instruct model has been fine-tuned with approximately 3,760,000 Thai instruction-completion pairs derived from human-annotated instructions, FLAN-style automatic data construction, and synthetic samples. This multilingual model supports both Thai and English languages.
  • WangchanLionv1 – WangchanLion 7B is a multilingual instruction-following model developed by PyThaiNLP and the VISTEC-depa AI Research Institute of Thailand. Fine-tuned on SEA-LION-v1-7B, it incorporates approximately 500,000 samples from open-source, commercially permissible datasets, with a focus on Thai and English languages.

Conclusion

By leveraging SEA-LION’s architecture, these localized models provide AI solutions that align more closely with native language requirements. As SEA-LION continues to evolve, more localized versions are expected to emerge, further expanding the reach and effectiveness of AI in Southeast Asian languages.

About the team

VISTEC-depa Artificial Intelligence Research Institute of Thailand is a collaborative effort between the Vidyasirimedhi Institute of Science and Technology (VISTEC) and the Digital Economy Promotion Agency (depa). Its mission is to advance artificial intelligence (AI) technology, equip Thai industry with advanced capabilities, and promote the widespread adoption of AI in Thailand.