Open Foundation Model
for Korea and Beyond
Lotus is a long-term research initiative to build an open, sovereign, and reproducible foundation model ecosystem based on open data, open models, open training code, and open evaluation.
Building independent AI infrastructure
Foundation models are becoming core infrastructure for national competitiveness, industrial productivity, and future AI services. Lotus aims to reduce dependency on external AI platforms and establish a sustainable open AI ecosystem for Korean language, knowledge, and industry-specific intelligence.
From 1B to 123B parameters
Lotus follows an incremental scaling strategy. Each phase expands model capacity, training tokens, Korean language capability, reasoning quality, coding performance, and domain applicability.
High-quality corpus and reproducible data pipeline
Lotus is designed around a large-scale Korean and English training corpus. Raw data is processed through collection, cleaning, validation, deduplication, quality scoring, and publishable dataset packaging.
Web, Wikipedia, Wikisource, news, literature, documents
HTML removal, spam filtering, broken document removal
Language checks, length checks, deduplication, scoring
JSONL, Parquet, metadata, Hugging Face dataset release
Llama-based decoder-only transformer
Lotus adopts a modern decoder-only transformer architecture to reduce training risk and maximize compatibility with the open source LLM ecosystem. The stack is optimized for scalable pretraining, efficient inference, and reproducible release.
From base model to domain expert models
Lotus evolves through four model layers: Base, Instruct, Chat, and Domain Expert. This structure enables the project to move from general language modeling to practical Korean enterprise use cases.
Next-token prediction
Supervised fine-tuning
DPO / ORPO / SimPO
Domain-adaptive continued pretraining
Specialized intelligence for real-world industries
On top of the foundation model, Lotus will support domain-specific expert models for legal, public, software, IT, semiconductor, and financial applications.
Open models, open code, open evaluation
Lotus aims to release datasets, tokenizer, training code, preprocessing scripts, SFT and preference alignment code, checkpoints, quantized models, and benchmark results through open platforms such as Hugging Face and GitHub.