How do you add AI to a Symfony application?
You add AI to a Symfony application the same way you add any external service: a configured HTTP client, a service class with a clear interface, background processing for slow work, and the usual security, logging and tests. Symfony already has every building block. HttpClient talks to the OpenAI or Claude API, Messenger runs long calls in workers, StreamedResponse streams answers, Doctrine DBAL queries pgvector, and the RateLimiter component caps usage per user.
Symfony is our main back-end stack, and this guide reflects how we build AI features into Symfony products. If you want a team to do it in your codebase, see our Symfony development services or, for older or framework-less code, our PHP development services.
The plan in short:
- A scoped HttpClient per provider, with keys from environment variables.
- One service interface for "ask the model", so providers can change.
- Messenger messages and handlers for every non-interactive call.
- A streamed endpoint for chat.
- pgvector for retrieval over your data, filtered by tenant and permissions.
- Rate limits per user and tenant, and a log table for tokens and cost.
Which AI features fit a Symfony product?
The AI features that fit a Symfony product best are the ones that sit next to business rules you already have: they read your entities, respect your voters and write results back through your services. Symfony products are often back offices, ERPs, marketplaces and SaaS platforms with many roles, which is exactly where these features save the most time.
| Feature | How it runs in Symfony | Typical first version |
|---|---|---|
| Summaries and drafts | Messenger handler, result stored on the entity | Ticket, order or report summaries reviewed by staff |
| Classification and routing | Messenger handler on create or import | Inbound requests tagged and assigned to a queue |
| Document extraction | Upload, then a handler that validates structured output | Invoice or form fields extracted into a review screen |
| Questions over your data | Streamed endpoint with retrieval from pgvector | An assistant over help articles or internal policies |
| Agent for one workflow | Tools that call your existing services, with approval steps | Preparing a weekly report or triaging a shared inbox |
Start with one of these, measured on real examples, before building a general assistant. Each row reuses the same plumbing described below, so the second feature is much cheaper than the first.
How do you call the OpenAI or Claude API with Symfony HttpClient?
Define a scoped client for each provider in configuration, so the base URL, headers, timeout and retry policy live in one place, then inject it by name.
# config/packages/framework.yaml
framework:
http_client:
scoped_clients:
anthropic.client:
base_uri: 'https://api.anthropic.com'
headers:
x-api-key: '%env(ANTHROPIC_API_KEY)%'
anthropic-version: '2023-06-01'
timeout: 120
retry_failed:
max_retries: 2
openai.client:
base_uri: 'https://api.openai.com'
auth_bearer: '%env(OPENAI_API_KEY)%'
timeout: 120
retry_failed:
max_retries: 2
retry_failed retries rate limits and server errors with backoff. Symfony autowires a scoped client by its camel-cased name, so HttpClientInterface $anthropicClient receives the first one.
namespace App\Ai;
use Symfony\Component\DependencyInjection\Attribute\Autowire;
use Symfony\Contracts\HttpClient\HttpClientInterface;
final class ClaudeModel implements LanguageModel
{
public function __construct(
private HttpClientInterface $anthropicClient,
#[Autowire(env: 'ANTHROPIC_MODEL')] private string $model,
) {}
public function complete(string $system, string $user, int $maxTokens = 4096): ModelResult
{
$data = $this->anthropicClient->request('POST', '/v1/messages', [
'json' => [
'model' => $this->model,
'max_tokens' => $maxTokens,
'system' => $system,
'messages' => [['role' => 'user', 'content' => $user]],
],
])->toArray(); // throws on 4xx and 5xx
$text = implode('', array_column(
array_filter($data['content'], fn (array $b) => $b['type'] === 'text'),
'text'
));
return new ModelResult($text, $data['usage']['input_tokens'], $data['usage']['output_tokens']);
}
}
An OpenAiModel class implements the same LanguageModel interface with a POST to /v1/responses. Bind the interface to one implementation in services.yaml, or pick per feature with a small router. That interface is what keeps a later provider switch, or an A/B test between models, a configuration change. Our Claude API vs OpenAI API comparison explains when each provider fits.
Anthropic publishes an official PHP SDK, and community PHP clients exist for OpenAI. They are fine choices; with Symfony's HttpClient you often do not need them for one or two features, and you keep Symfony's profiler, retry and mocking tools.
How do you run LLM calls asynchronously with Symfony Messenger?
Dispatch a message for every AI task the user does not watch live, and let a dedicated worker handle it with a retry strategy. This keeps web requests fast and turns provider rate limits into delays instead of errors.
// src/Message/SummarizeDocument.php
final class SummarizeDocument
{
public function __construct(public readonly int $documentId) {}
}
// src/MessageHandler/SummarizeDocumentHandler.php
#[AsMessageHandler]
final class SummarizeDocumentHandler
{
public function __construct(
private LanguageModel $model,
private DocumentRepository $documents,
private PromptLibrary $prompts,
private EntityManagerInterface $em,
) {}
public function __invoke(SummarizeDocument $message): void
{
$document = $this->documents->find($message->documentId)
?? throw new UnrecoverableMessageHandlingException('Document deleted');
$result = $this->model->complete(
$this->prompts->get('document-summary', version: 4),
$document->getText(),
maxTokens: 1024,
);
$document->setSummary($result->text);
$this->em->flush();
}
}
# config/packages/messenger.yaml
framework:
messenger:
transports:
ai:
dsn: '%env(MESSENGER_TRANSPORT_DSN)%'
options: { queue_name: ai }
retry_strategy:
max_retries: 3
delay: 5000
multiplier: 4
routing:
App\Message\SummarizeDocument: ai
Run the ai transport with its own workers (messenger:consume ai --time-limit=3600) so a backlog of AI work never delays emails or payments. Throw UnrecoverableMessageHandlingException for errors a retry cannot fix, such as a deleted record or a 400 response, so the message goes to the failure transport instead of retrying for nothing.
How do you stream AI responses in Symfony?
Use EventSourceHttpClient to read the provider's Server-Sent Events and a StreamedResponse to forward only the text to the browser. The first words appear in about a second.
use Symfony\Component\HttpClient\Chunk\ServerSentEvent;
use Symfony\Component\HttpClient\EventSourceHttpClient;
use Symfony\Component\HttpFoundation\StreamedResponse;
#[Route('/assistant/stream', methods: ['POST'])]
#[IsGranted('ROLE_USER')]
public function stream(Request $request, HttpClientInterface $anthropicClient): StreamedResponse
{
$question = (string) $request->getPayload()->get('q');
$model = $this->getParameter('app.anthropic_model');
$client = new EventSourceHttpClient($anthropicClient);
return new StreamedResponse(function () use ($client, $question, $model) {
$source = $client->request('POST', '/v1/messages', ['json' => [
'model' => $model,
'max_tokens' => 4096,
'stream' => true,
'messages' => [['role' => 'user', 'content' => $question]],
]]);
foreach ($client->stream($source) as $chunk) {
if (!$chunk instanceof ServerSentEvent) {
continue;
}
$event = json_decode($chunk->getData(), true);
if (($event['type'] ?? null) === 'content_block_delta'
&& ($event['delta']['type'] ?? null) === 'text_delta') {
echo 'data: '.json_encode(['text' => $event['delta']['text']])."\n\n";
flush();
}
}
echo "event: done\ndata: {}\n\n";
flush();
}, 200, [
'Content-Type' => 'text/event-stream',
'Cache-Control' => 'no-cache',
'X-Accel-Buffering' => 'no',
]);
}
Recent Symfony versions add a dedicated response class for event streams; check the docs for your version. On PHP-FPM, each open stream holds a worker for its whole duration, so size the pool for concurrent chats, or run the app on FrankenPHP with the Runtime component. If you already use Mercure for real-time updates, as we did in the custom ERP for a US aircraft service company, a worker can also publish partial answers to a Mercure topic instead of holding an HTTP connection open.
How do you store embeddings with Doctrine and pgvector?
Add the pgvector extension to PostgreSQL, create the vector column and index in a Doctrine migration with raw SQL, and query through the DBAL connection with tenant filters in the same statement.
// in a Doctrine migration
$this->addSql('CREATE EXTENSION IF NOT EXISTS vector');
$this->addSql('ALTER TABLE chunk ADD embedding vector(1536)'); // match your embedding model
$this->addSql('CREATE INDEX chunk_embedding_idx ON chunk USING hnsw (embedding vector_cosine_ops)');
final class ChunkSearch
{
public function __construct(private Connection $connection) {}
/** @param float[] $embedding */
public function nearest(int $tenantId, array $embedding, int $limit = 8): array
{
return $this->connection->fetchAllAssociative(
'SELECT id, document_id, content
FROM chunk
WHERE tenant_id = :tenant
ORDER BY embedding <=> CAST(:q AS vector)
LIMIT '.(int) $limit,
['tenant' => $tenantId, 'q' => '['.implode(',', $embedding).']']
);
}
}
Embeddings come from an embedding model such as OpenAI's embeddings endpoint; Anthropic does not offer its own embedding models, so Claude-based projects pair it with OpenAI or another provider. Generate embeddings in a Messenger handler when documents change, and keep a hash of the content so unchanged text is not re-embedded. Doctrine entities do not need to know about the vector column at all. For chunking, hybrid search, reranking and evaluation, read our RAG implementation guide.
How do you rate-limit AI features in Symfony?
Use the RateLimiter component with a limiter per user and per tenant, and check it before every model call. A model in a loop, or a user who scripts your chat, should hit a limit, not your invoice.
# config/packages/rate_limiter.yaml
framework:
rate_limiter:
ai_per_user:
policy: 'token_bucket'
limit: 30
rate: { interval: '10 minutes', amount: 30 }
public function __construct(private RateLimiterFactory $aiPerUserLimiter) {}
public function ask(User $user, string $question): ModelResult
{
$limit = $this->aiPerUserLimiter->create((string) $user->getId())->consume(1);
if (!$limit->isAccepted()) {
throw new TooManyRequestsHttpException($limit->getRetryAfter()->getTimestamp() - time());
}
// ... call the model
}
Combine limits with the cost levers that apply to any stack: prompt caching for long stable instructions, a smaller model for simple steps, batch APIs for nightly work, and a log table with feature, prompt_version, model, token counts, latency and cost per call. A monthly budget per tenant can be checked against that table before a call.
How do you test and secure AI features in Symfony?
Test your code with MockHttpClient and recorded provider responses, and test quality separately with an evaluation set. Treat everything the model reads as untrusted input.
$client = new MockHttpClient([
new JsonMockResponse([
'content' => [['type' => 'text', 'text' => 'Invoice 1042 is overdue by 12 days.']],
'usage' => ['input_tokens' => 300, 'output_tokens' => 14],
]),
]);
Register the mock as the scoped client in the test environment, then assert on what your handler saved, what it logged and how it handles a 429 or a timeout. For quality, keep 50 to 200 real inputs with expected outputs and run them on every prompt or model change.
Security rules we apply on every Symfony AI feature:
- Keys only on the server, injected from environment variables or Symfony secrets.
- Voters before retrieval. Only documents the user can read are searched.
- Untrusted text stays data. User input and retrieved documents are clearly delimited in the prompt and never treated as instructions.
- Validate structured output with the Validator component before it reaches the database.
- Approval for actions. Anything that sends, pays or deletes needs a human confirmation.
What do we recommend for Symfony teams?
Our verdict for most Symfony products: use scoped HttpClients behind your own LanguageModel interface, Messenger for every non-interactive call, a streamed endpoint for chat, pgvector through DBAL, and the rate limiter from day one. Watch Symfony's AI initiative and adopt its components when they are stable and solve a problem you actually have. This setup uses only components your team already knows, which is the main reason it holds up in production.
Our own AI SaaS, AI Resume Master, runs on PHP and Symfony with React and Python and the OpenAI API; it was built in about three months and reached 50,000 monthly active users. If your app is on Laravel, the same ideas are in our guide on adding OpenAI or Claude to a Laravel app.
Next step
If you want a Symfony team to add an AI feature to your product, see our Symfony development services. AI features in existing products start from $10,000 with us, with a fixed quote after a short scoping call. Contact us with the feature you have in mind and the data it needs.