What an AI audit is, and what it is not
An AI audit is an engineering review of one existing software product with AI in mind. It answers three practical questions: where AI features would genuinely help this product, whether its data and codebase can support them, and, if the product already uses AI, how well that AI works on accuracy, security and cost.
It is not a company-wide AI strategy. If you want to compare AI ideas across sales, support, finance and operations, that is AI consulting. An AI audit starts from the code and the data of a product you already run, and ends with findings an engineering team can act on next sprint.
We run three kinds of AI audit, often combined:
| Audit type | For whom | Main question |
|---|---|---|
| AI fit audit | Products without AI features yet | Where should AI go in this product, and what will it take? |
| LLM feature audit | Products with AI features in production | Is our AI accurate, safe and affordable, and how do we fix it if not? |
| AI-generated code audit | Products built largely with AI coding tools | Is this code ready for real users, real data and growth? |
AI fit audit: where AI belongs in your product
An AI fit audit maps your product's workflows, data and architecture to concrete AI features, and tells you which ones are worth building first.
What we review:
- User workflows where people read, write, search, classify or copy data between screens. These are the usual candidates for drafting, summarizing, extraction and search features.
- Data: where it lives, its format and quality, personal and confidential fields, and whether there is enough of it, in the right shape, for retrieval or evaluation.
- Architecture: where an AI service would sit, how it would call your APIs, how permissions would carry through, and what would need refactoring first.
- Model options: prompting, retrieval (RAG), fine-tuning or an agent with tools, and which model families to test.
What you get: a ranked list of AI features with value, feasibility, risk and estimated running cost per request; data readiness notes per feature; an architecture sketch; and a cost range for building the top candidates.
LLM feature audit: checking the AI you already run
Many products shipped an LLM feature quickly and now see one of three symptoms: answers users do not trust, a model bill that grows faster than revenue, or nobody knowing whether the last prompt change made things better or worse. An LLM feature audit finds the causes.
| Area | What we check |
|---|---|
| Quality | Is there an evaluation set? We build a small one from your real inputs and measure current accuracy and failure patterns. |
| Grounding | Do factual answers come from your data with sources, or from model memory? How good is retrieval? |
| Prompts | Are prompts versioned, tested and separated from user input? |
| Security | Prompt injection paths, data sent to the provider, permission checks on tool calls, key handling. |
| Cost | Cost per request and per user, caching, model routing, context size, limits against abuse and loops. |
| Reliability | Timeouts, retries, fallbacks between providers, behaviour when the model API is slow or down. |
| Observability | Logs of prompts, outputs, token counts and errors, and whether anyone looks at them. |
The output is a prioritized list of fixes with expected effect, for example "add prompt caching to the three largest calls" or "move these answers onto retrieval with citations", plus the evaluation set we built, which you keep using after the audit.
AI-generated code audit: before real users arrive
Products built quickly with AI coding assistants can reach a convincing demo in days. The gaps usually show up later: permissions checked in the interface but not in the API, inputs not validated, secrets in the repository, no tests, duplicated logic, and database queries that collapse under real data.
We use AI coding agents in our own delivery every day, under senior engineers who own architecture and code review, so we know the typical failure patterns of AI-written code. The audit covers:
- Security: authentication, authorization on every endpoint, input validation, secrets, dependency risks.
- Data model and queries: integrity, migrations, performance with realistic volumes.
- Architecture: separation of concerns, duplication, how hard the next feature will be.
- Tests and delivery: coverage of critical paths, CI, deployment and rollback.
The report separates must-fix-before-launch issues from things that can wait, with an effort estimate for each.
How an AI audit runs
- 30-minute call. We agree which audit types apply, what access is needed and the fixed price.
- Access and kickoff. NDA, read-only repository access, data samples, logs and a walkthrough with your lead engineer.
- Review (1 to 2 weeks). Code and data review, an evaluation set for existing AI features, cost analysis from your logs and bills, and quick tests of two or three models where AI fit is in scope.
- Report and walkthrough. A written report and a call with your team to go through findings and priorities.
Typical length is 1 to 3 weeks, depending on the size of the codebase and the number of AI features already in production.
What you receive
- A written audit report in English: findings, severity, evidence and recommended fixes.
- A prioritized plan: what to do first, what can wait, and what not to build.
- Cost estimates: build effort for recommended features or fixes, and running cost per request and per month for AI features.
- Risk register: security, privacy, quality and vendor risks, each with a mitigation.
- Reusable assets: the evaluation set and any test scripts we wrote during the audit.
Who runs the audit
Senior engineers who build and run AI products, not reviewers working from a checklist. We built AI Resume Master, our own AI SaaS on the OpenAI API, in about three months and took it to 50,000 monthly active users, so we have met the cost, quality and abuse problems that appear only at scale. For a US grief-tech startup we built the AI Grief Companion, with ingestion of ten chat-export formats, LoRA fine-tuning, RAG memory and per-message AES-256-GCM encryption at ingest, which is the level of care we look for when we audit sensitive data flows.
Our main back-end stack is PHP and Symfony, with React and Next.js on the front end, and we review Node, Python and mobile codebases as well. We work with OpenAI, Anthropic Claude and open models, and we build and run MCP servers in production for our own systems.
What an AI audit costs
We quote a fixed price after the first call, based on codebase size, the audit types in scope and the number of AI features already running. If you continue with us, the audit counts as the scoping phase of the build, so you do not pay for the same analysis twice. AI features for an existing product start from $10,000; a first feature or a proof of concept can run as one fixed-price AI Sprint for $10,000 in 4 weeks.
Lytvynov Production was founded in 2020 in Dnipro, Ukraine, and is rated 5.0 on Upwork with 100% Job Success.
Next step
Tell us what the product does, what stack it runs on and whether it already uses AI. In a 30-minute call we will say which audit fits, what access we need and what it will cost. Contact us to book the call.