Buyer Guides · 6 MIN
IP and data terms to settle before an AI project starts
A buyer checklist of the IP, data, and model-provider questions to settle before an AI project starts. A discussion guide for your counsel, not legal advice.
Before an AI project starts, settle six things in writing: who owns the code, who owns the prompts and eval sets, what the partner may do with your data, which model providers may see it, how outputs and any trained artifacts are owned, and what happens at exit. This guide is a checklist of questions to bring to the table. It is not legal advice, and your counsel should review the actual terms.
- Settle ownership of code, prompts, evaluation sets, and any fine-tuned artifacts explicitly.
- Define what data the partner receives, how it is stored, and when it is deleted.
- Name every model provider that will process your data and the terms that apply.
- Decide whether your data may ever be used to train a model, and write the answer down.
- Plan the exit: handover of repositories, documentation, and access.
- Have your own counsel review everything. Nactore designs engagements so you own what we build for you.
Why do AI projects need more than a standard software contract?
A standard development agreement covers code ownership and confidentiality. AI projects add layers a normal contract may not address. Your data flows through third-party model providers. The project may produce prompts, evaluation sets, and tuned models that are valuable on their own. And questions arise about whether data can be retained or used to improve a model.
None of this is exotic. It simply needs to be stated, because defaults vary by provider and by supplier, and assumptions are where disputes begin.
Who should own what the project produces?
List each artifact and decide ownership explicitly.
| Artifact | Why it matters | Question to settle |
|---|---|---|
| Application code | The core deliverable | Is it assigned to you on payment or delivery? |
| Prompts and orchestration logic | Often the real intellectual content | Do you own them, and can you modify them freely? |
| Evaluation sets and labels | Built from your data and expertise | Are they yours, stored in your repositories? |
| Fine-tuned models or adapters | Trained on your data | Who owns the artifact and the right to use it? |
| Generated outputs | The system's day-to-day product | Are they yours without restriction? |
| Supplier's pre-existing tools | Used to speed the work | What license do you receive to keep running them? |
The last row deserves attention. Partners often bring reusable internal tooling. That is normal and can speed delivery, but you need a license that lets you operate the system after the engagement ends.
What should the data terms cover?
Write down the full lifecycle of your data. Your counsel will want to see these points addressed.
- What data is shared. Describe categories and sensitivity, and prefer sharing the minimum needed.
- Where it is stored. Include location, access controls, and who on the partner side can see it.
- How long it is kept. Set a retention period and a deletion commitment with confirmation at the end.
- Personal data handling. If personal data is involved, identify the roles and the regulatory regime that applies, such as GDPR or UK GDPR for European and UK data subjects. Counsel should decide the right structure.
- Security measures. Ask for a description of controls and how incidents are reported.
Which model providers will see your data?
Every API call to a hosted model sends your data to a provider. List each provider and check its published terms for your plan, since they differ between consumer products and business or API offerings.
Questions to ask for each provider include whether inputs are retained, for how long, whether they are used for training by default, and what opt-out or zero-retention options exist. Read the provider's own documentation, for example OpenAI's data controls documentation and Anthropic's commercial terms and privacy center, rather than relying on a summary from a seller.
If the data is too sensitive for any third-party API, ask about options such as private deployments or self-hosted models, and what they would change in cost and quality. For a related engineering choice, see choosing an LLM model for production.
Ask the partner for a one-page data-flow diagram showing every system and provider that touches your data. If they cannot draw it, they have not thought it through.
Can your data be used to train models?
Decide this deliberately and write the answer down. Most buyers want a clear "no" for their proprietary or customer data, both for the partner's own purposes and for upstream providers. Some buyers do want a tuned model and accept training on their data for their own use only.
Either answer can be right. The important part is that the contract, the provider settings, and the actual implementation all agree. Settings often default one way, so verify them rather than assuming.
What happens at the end of the engagement?
Plan the exit before you start, while everyone is on good terms. A reasonable exit clause covers the following.
- Handover. All code, prompts, configuration, eval sets, and documentation delivered into repositories you control.
- Access. Transfer or removal of accounts, API keys, and infrastructure, with keys rotated.
- Data deletion. Confirmed deletion of your data from partner systems, with a date.
- Transition support. A defined period of help for your team to take over.
How do you get this reviewed efficiently?
Bring your counsel a short summary of the above questions before the contract draft arrives. It speeds review and surfaces your priorities. Also involve your security or compliance lead early, since their sign-off on providers and data paths often gates the project timeline, as covered in how to scope an AI pilot.
Frequently asked questions
Does the client usually own the code a partner writes?
It is common for the client to own custom code, but it depends entirely on the contract. Confirm it explicitly, along with the license to any reusable partner tools. Your counsel should review the wording.
Is it safe to send company data to a hosted model API?
It depends on the data, the provider's terms for your plan, and your regulatory obligations. Review the provider's published terms and involve your security and legal teams before sending sensitive data.
Who owns a fine-tuned model?
It should be stated in the contract. Decide whether the artifact is yours, who may use it, and where it is stored, then confirm with counsel.
Should we worry about our data being used for training?
Check the provider's terms and settings for your plan, and write your requirement into the agreement. Verify the configuration in practice, not only on paper.
Write it down before you build
Clear terms on ownership, data, providers, and exit protect both sides and rarely slow a project. Settle them first, then build.
Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.