Domain-adaptive pretraining (DAPT)
Beginner
Giving a general AI model extra study on one field’s documents, such as chip-design manuals and code, so it knows that field better.
Novice
Continuing to train a general language model on one field’s documents and code, so it learns that field’s vocabulary and facts, before teaching it to follow instructions.
Expert
Costs a small fraction of the original pretraining compute but needs a large, clean domain corpus, which in chip design is proprietary. Applied to a model already tuned for chat, it can wreck the model’s instruction following, so the instruction tuning is done afterward.