GLOSSARY · AI SECURITY

Model extraction

Model extraction is an attack that reconstructs a proprietary AI model, or the sensitive data behind it, by systematically querying it and studying the responses.

Attackers can approximate a model’s behavior, steal fine-tuned intellectual property, or use related inference attacks to test whether specific records were in the training data.