Using an AI agent on licensed data without risking your campus's access
Most data licenses forbid putting the content into AI systems. An agent can still write and repair the whole pipeline if the licensed rows never enter its context.
At a major conference in 2026, 22.5% of reviewers who were told not to use an LLM said they used one anyway. The rules are already losing. You can get the speed without betting your campus’s access on it.
Two libraries say no to chatbots and yes to built-in AI tools
The answers come from the published guidance of two university libraries, one in Canada and one in the United States. Your own library’s answer is the one that binds you.
| What you want to do | Usually allowed? | Why |
|---|---|---|
| Paste licensed data or a licensed article into a chatbot | No | Both libraries say licenses forbid it |
| Do the same in an AI tool your university pays for | No | American University says the restriction covers “all AI tools,” including university-licensed ones |
| Use the AI search or assistant built into a library database | Yes | American University says embedded tools “are designed to work within the licensing framework” |
| Use open access or public domain material with any AI tool | Yes | Not governed by the license |
| Run your own code over licensed data on your own machine | Depends | Waterloo treats text and data mining as a separate question from generative AI. Check the license |
| Run a local model over licensed data on your own machine | Unsettled | Nothing leaves your machine, but “using the content in AI systems” may still cover it. Get an answer in writing |
Waterloo also notes where the answer is recorded: when a publisher permits generative AI use, the library says it will be noted in its catalogue. Look for the equivalent at your school before you email anyone.
Keep the rows out of the model’s context
The restriction is on the content reaching the model. Most of what an agent does for you does not need the content.
- Give the agent the shape of the data, not the data. Table names, column names, types, and a few rows you made up. That is enough for it to write the query and the cleaning code.
- Let it write code that you run. The agent writes
pull.py. You run it. The licensed rows go from the vendor to your disk and never enter the model’s context. Our data pull guide is built this way. - Return summaries, not rows. Have scripts print row counts, date ranges and test results. The agent can debug from a manifest.
- Keep outputs that contain licensed values out of the chat. A regression table of coefficients is yours. A printout of raw vendor fields is not.
- Write the boundary into the instruction file. One line in the repository that says which folders the agent must never read or print. It reads that file on every run.
Five questions for your library
- Does our license for this database permit use with generative AI? Where is that recorded?
- Does the answer change for a model that runs on my own machine?
- Does it change for a cloud service where the model provider cannot see prompts?
- Is text and data mining covered separately?
- Who signs off if I need an exception for a project?
Ask in writing and keep the reply. If the answer is no, ask what it would take at the next renewal, and tell your department head you asked. Licenses are renegotiated. Schools that ask for agent use get it sooner than schools that stay quiet.
Keep a note of which tool touched which dataset
For each project, keep a short note: which datasets, under which license terms, which tools touched what. A referee or a vendor may ask one day, and you will want the answer in one paragraph.
References
- Use of library resources with AI, University of Waterloo Libraries uwaterloo.ca
- May I upload library electronic resources into an AI tool or LLM?, American University Library answers.library.american.edu
- Amazon Bedrock, data protection docs.aws.amazon.com