
Personal data ends up in AI pipelines by accident more than by design. A support ticket contains a name and an address, a document carries a date of birth, a transcript includes a card number the customer read out. All of it flows into prompts, and from there into logs, caches and third-party systems. Handling personal data in an AI pipeline is mostly about knowing where it accumulates, because the copies nobody planned are the ones with no controls on them. Prompts and logs are the quiet problem Careful teams protect their database and then log every prompt in full
