Skip to main content
AiGORA
Back to Q&A
Text
4

What happens to data that I (involuntarily) provide to a GenAI tool?

The other day, I asked CoPilot something related to my career progression, and the output addressed me by name and knew what type of contract I am currently employed on, which I found quite scary. I know that we aren't supposed to upload student names and other sensitive data to GenAI, so the university seems to have some data privacy concerns there, but at the same time, if the university-provided (and, I am assuming, approved) tool can access my files and emails and browser, I would assume it can access all that data anyway. What happens to this data, what happens to the questions I ask, the input I provide? Can anyone see all of it directly, does it somehow become part of the training data, ...? I am concerned!

Anonymous · 11 Sept 2026

Responses from the team

2 perspectives from the community

  • Rachel Forsyth's profile photo

    Rachel Forsyth

    What's happening here is that your employer has an institutional contract with Microsoft to give you access to Copilot. If your name and role are used, it's because those are considered part of institutional knowledge as part of your Microsoft account. Alongside the contract, there is a data processing agreement which controls where that information can be shared - which is usually only within the institution. The agreement should also say that your name/ID are unlinked from the queries you make, so cannot be sent outside the university (since the ID needs to be linked back to the response Copilot makes, you have to wonder if it is actually possible to recreate that link), and also that your queries will not be used for future Large Language Model training.

    Some types of Copilot contract also allow Copilot access to all of your Microsoft managed files: Outlook for email, OneDrive for documents, OneNote for notes, and so on. You should know if you have this access, since it costs extra and someone will have checked up whether you want it. So a question to Copilot might bring a response which references files you have saved. If you are wondering why anyone would want this, it could be useful for helping you to pull together relevant information. Requests like "list all the files which are related to X project" might bring up some (ahem) misfiled items or meeting notes you have forgotten, or help you combine information from different types of files (Excel and Powerpoint, for instance). But if you have this option and find it creepy, ask someone to cancel your subscription: every little helps the university budget! Whilst personal information you hold in a limited way for the purposes of your job should be secure within this kind of data agreement, you can move files you definitely don't want the GenAI product to have access to into a different location which Copilot does not have access to (e.g. the hard drive of your computer). If you aren't sure what is allowed, check with your organisation's data controller.

    It's worth noting that this is a question where digital inequalities rear their heads, alongside very reasonable concerns about data privacy. Students may not have access to this 'safe' way of using a GenAI product with a reliable data processing agreement. Just using the free versions of products may put them at risk of their chats being exposed in security breaches, used for training, and they may, intentionally or not, also be sharing copyright, personal or sensitive data more widely by uploading to a GenAI product. They may also not have access to the most recent version of the products, but that isn't the focus of your question.

    Read all posts that Rachel has responded to →

  • Mirjam Glessmer's profile photo

    Mirjam Glessmer

    Here is an interesting case of GenAI being used to solve the Navier-Stokes problem, one of the "Millenium Prize Problems", where there is at least the suspicion that storing ongoing work in a model might have contributed to its training -- without the knowledge of the actual humans who stored their ongoing work there:

    https://www.theguardian.com/science/2026/sep/08/openai-claims-to-have-solved-maths-problem-that-stumped-humans-for-decades

    I don't know what kind of contracts they had with the GenAI providers, but it's at the least interesting to be aware that this kind of thing might happen, and concern about data privacy seems to not be such an unreasonable reaction.

    Read all posts that Mirjam has responded to →

Comments

Share your thoughts — comments are reviewed before they appear

No comments yet. Be the first!