Acceptable Use Policy Generator for AI companies
Written for training data, prompts and outputs, the EU AI Act and the questions enterprise buyers actually ask.
For an AI product the acceptable use policy is a safety control, not just a legal one. It is what a trust and safety team enforces against, what an enterprise reviewer reads to assess your maturity, and what a model provider requires you to pass through to your own users.
AI companies face a disclosure problem no template handles: the same data can be an input, a training corpus, an output and a log, and each of those has a different legal character. A prompt containing personal data is processing; retaining it is storage; training on it is a new purpose; and generating an output about a person is processing again, with an accuracy obligation attached.
Regulators have moved fast here. Data protection authorities have opened and settled investigations into training data provenance, lawful basis for web scraping, the accuracy of model outputs about real people, and the adequacy of opt-outs. The EU AI Act adds a transparency layer of its own: people must be told when they are interacting with an AI system, synthetic content needs marking, and general-purpose model providers face documentation duties.
Commercially, the questions that actually decide deals are narrow and specific: do you train on customer data, how long do you retain prompts, which model providers are in the chain, can data be processed in a specified region, and can the customer opt out of everything above. A policy that answers those five questions plainly beats one that is comprehensive but evasive.
What a acceptable use policy for an AI company has to cover
Prohibited use categories that map to real harm rather than to vague reputational risk
Restrictions passed through from your model providers’ own usage policies
Automated enforcement, human review, and how a decision is appealed
Rate limits, scraping and circumvention of safety systems
Reporting route for harmful outputs, with a stated response process
How an AI company actually moves personal data
Prompts and inputs
Users paste anything into a prompt box, including personal data about third parties. Retention of prompts, and who inside the company can read them, is the first question in every review.
Outputs and generated content
Outputs about identifiable people are personal data. That brings accuracy, rectification and erasure obligations to content the model produced rather than content anyone stored.
Training and fine-tuning corpora
Scraped data, licensed datasets, and customer content if you use it. Provenance, basis and opt-out routes all need documenting, and the answer materially affects enterprise sales.
Human review and evaluation
Rating pipelines and red-teaming expose staff or contractors to user content. It is legitimate processing but it needs disclosing, because users assume nobody reads their prompts.
Model provider sub-processing
Calls to OpenAI, Anthropic, Google or an inference host move customer content to a third party with its own retention window and its own training position.
Telemetry, logs and abuse monitoring
Safety and abuse systems retain content longer than the product does, which is defensible but has to be described rather than discovered.
Third parties the draft will ask you about
OpenAI · Anthropic · Google Cloud Vertex · AWS Bedrock · Pinecone or Weaviate · Modal or Replicate · Scale or Surge for human evaluation · Datadog
The rules that apply
Lawful basis for training data
Scraped and licensed corpora containing personal data need a basis, and legitimate interests requires a documented balancing test that engages with the individuals’ expectations.
GDPR Article 22 and automated decisions
Where an AI system makes decisions with legal or similarly significant effects, the individual has rights to information, human intervention and challenge.
EU AI Act transparency duties
Disclosure when a person interacts with an AI system, marking of synthetic content, and documentation obligations for general-purpose models, phasing in on a set timetable.
Accuracy and rectification
A model that produces false statements about a real person engages the accuracy principle, and "the model generated it" is not a defence regulators have accepted.
Model provider chain
If you route customer content to a third-party model, that provider is a sub-processor with its own retention and training terms that flow through to your customers.
What the generated acceptable use policy contains
Prohibited content and conduct
The categories you will not host, written specifically enough to enforce and narrowly enough to defend.
Technical restrictions
Rate limits, scraping, load testing, resource abuse and circumvention of quotas.
Security research boundaries
What testing is permitted, how to report a finding, and what will not be treated as an attack.
Reporting and moderation
How abuse is reported, what happens next, and the timescales you work to.
Enforcement ladder
Warning, throttle, suspension, termination - with the emergency route for genuinely urgent harm.
Appeals
How someone contests a decision, which several online-safety regimes now require rather than merely encourage.
The AI compliance document set
Write the training position first
One paragraph, in plain language: whether customer content trains models, whether there is an opt-out, and what happens to data already used.
Set and publish retention windows
For prompts, outputs, logs and abuse-monitoring copies - which are usually different from each other.
Document training data provenance
Sources, licences, and the legitimate interests assessment for anything scraped.
Build a rectification route for outputs
A way for someone to report a false generated statement about themselves, and a documented response.
Publish the model provider chain
Which providers process customer content, in which regions, with what retention and training terms.
Add AI Act transparency where it applies
Interaction disclosure, synthetic content marking, and the documentation duties for general-purpose models.
Where this usually goes wrong
Silence on training
If the policy does not say whether customer content trains models, buyers assume the worst answer and regulators treat the omission as a transparency failure.
An opt-out that only covers future data
If a user opts out, the honest position on data already in a training set has to be stated - including where removal is not technically possible.
No accuracy or rectification route for outputs
People have complained to regulators about false generated statements. Having no process is worse than having an imperfect one.
Undisclosed human review
Users are consistently surprised that humans read prompts. Disclosure costs almost nothing; discovery costs a great deal.
Retention windows set by the model provider, not by you
Provider defaults - often thirty days for abuse monitoring - become your retention period unless you configure otherwise, and your policy should reflect what is actually configured.
Treating the AI Act as a future problem
The transparency obligations arrive on a schedule, and the documentation they require takes longer to assemble than the deadline suggests.
Frequently asked questions
Do I need to disclose that I train on user data?
Yes. Training is a distinct purpose from providing the service, it needs its own lawful basis, and both regulators and enterprise buyers treat silence as an answer. The disclosure should be specific enough to act on, not a general reservation of rights.
Are AI outputs personal data?
Where they relate to an identifiable person, yes - which means accuracy, rectification and erasure obligations apply to generated content, not just to stored content.
What does the EU AI Act require of my privacy documentation?
Transparency where people interact with an AI system, marking of synthetic content, and documentation for general-purpose models. It sits alongside GDPR rather than replacing it, and the obligations phase in on a published timetable.
Is a model provider a sub-processor?
If customer content is sent to it, yes. That brings a contractual chain, a transfer question, and an obligation to tell your customers who is in the chain and on what terms.
Can users ask to be removed from a training set?
They can ask, and you have to answer honestly. Where removal from an already-trained model is not technically feasible, say so and explain what you can do instead - exclusion from future training, output filtering, or deletion of the source record.
Do I need an acceptable use policy separate from my terms?
A separate AUP is easier to enforce and easier to update. It also gives moderation staff and automated systems a single reference to cite, which matters when a suspension is challenged.
Does an AUP help with platform liability?
It is part of the picture. Intermediary liability protections generally depend on acting on notice, and both the EU Digital Services Act and the UK Online Safety Act expect published rules, a reporting route and an appeals process.
How specific should prohibited-use lists be?
Specific enough that a moderator can apply it consistently, with a residual catch-all. Lists that are only catch-alls get challenged; lists that are only specific leave gaps.
Acceptable Use Policy Generator for AI companies
Answer a short questionnaire and get a draft written for an AI company. Free to start, no card required.
Generate your acceptable use policyOther documents an AI company needs
Each one is written for the same context, not a generic template.
The same document, by business type
Go deeper
PolicifyAI is a technology provider, not a law firm, and this page is not legal advice. Generated documents are a structured starting point that a qualified adviser should review before you publish or rely on them.