What does it take to add OpenAI or Claude to a Laravel app?

Adding OpenAI or Claude to a Laravel app means writing a service class that sends prompts to the provider's API from your server, then building the usual Laravel plumbing around it: queued jobs for slow work, a streamed response for chat, a table that logs every call, retrieval over your own data, rate limits and tests. The API call is ten lines. The plumbing is what makes it safe to ship.

This guide shows working code for each piece, using current Laravel conventions. It is written for Laravel developers and the CTOs who review their plans. If you would rather have a team do it inside your codebase, our Laravel development services cover AI features in existing Laravel apps, and our AI integration services describe the broader approach.

The short version:

  1. Keep API keys in config and call the API only from the back end.
  2. Put every provider call behind one service class.
  3. Run slow calls in queued jobs with timeouts and retries.
  4. Stream interactive answers with Server-Sent Events.
  5. Log prompts, versions, tokens and cost for every call.
  6. Add retrieval (RAG) with PostgreSQL and pgvector when answers depend on your data.
  7. Rate-limit per user and per tenant.
  8. Test with Http::fake, and measure quality with an evaluation set.

Should you use Laravel's HTTP client or an SDK package?

Use Laravel's HTTP client when you have one or two AI features and want full control with no extra dependency. Use an SDK package when you need typed responses, provider-specific features or several providers at once. Either way, hide it behind your own class.

Option Good for Watch out for
Laravel HTTP client (Http::) Few features, full control, easy Http::fake tests You maintain request and response shapes yourself
Anthropic's official PHP SDK Claude-heavy features, typed responses, streaming helpers Check the current SDK docs for method names; they change between versions
Community OpenAI packages for PHP and Laravel OpenAI features with a facade and test fakes Community-maintained, so check release activity
Multi-provider community packages Switching between OpenAI, Claude and others with one API Another abstraction layer to keep up to date

Start by adding keys to config/services.php. The model name lives in config too, so changing it is not a code change.

// config/services.php
'anthropic' => [
    'key' => env('ANTHROPIC_API_KEY'),
    'model' => env('ANTHROPIC_MODEL'),
],
'openai' => [
    'key' => env('OPENAI_API_KEY'),
    'model' => env('OPENAI_MODEL'),
    'embedding_model' => env('OPENAI_EMBEDDING_MODEL'),
],

A minimal Claude client with the HTTP client, calling the Messages API:

namespace App\Ai;

use Illuminate\Http\Client\ConnectionException;
use Illuminate\Support\Facades\Http;

final class ClaudeClient
{
    public function complete(string $system, string $user, int $maxTokens = 4096): array
    {
        $response = Http::withHeaders([
                'x-api-key' => config('services.anthropic.key'),
                'anthropic-version' => '2023-06-01',
            ])
            ->timeout(120)
            ->retry(2, 1000, fn ($e) => $e instanceof ConnectionException, throw: false)
            ->post('https://api.anthropic.com/v1/messages', [
                'model' => config('services.anthropic.model'),
                'max_tokens' => $maxTokens,
                'system' => $system,
                'messages' => [['role' => 'user', 'content' => $user]],
            ])
            ->throw();

        return [
            'text' => collect($response->json('content'))
                ->where('type', 'text')->pluck('text')->implode(''),
            'input_tokens' => $response->json('usage.input_tokens'),
            'output_tokens' => $response->json('usage.output_tokens'),
        ];
    }
}

The OpenAI equivalent uses the Responses API. The answer text sits inside output items of type message:

$response = Http::withToken(config('services.openai.key'))
    ->timeout(120)
    ->post('https://api.openai.com/v1/responses', [
        'model' => config('services.openai.model'),
        'instructions' => $system,
        'input' => $user,
    ])
    ->throw();

$text = collect($response->json('output'))
    ->where('type', 'message')
    ->flatMap(fn (array $item) => $item['content'])
    ->where('type', 'output_text')
    ->pluck('text')
    ->implode('');

Both responses report usage.input_tokens and usage.output_tokens, which is all you need for cost tracking. Not sure which provider to start with? Our comparison of the Claude API vs the OpenAI API covers the trade-offs.

How do you handle long AI calls with Laravel queues?

Put every AI call that the user does not watch live into a queued job: summaries, tagging, extraction, document processing, nightly reports. LLM calls can take from a few seconds to over a minute, and a web worker waiting on them is a worker not serving anyone else.

namespace App\Jobs;

use App\Ai\ClaudeClient;
use App\Models\Ticket;
use Illuminate\Contracts\Queue\ShouldQueue;
use Illuminate\Foundation\Queue\Queueable;

final class SummarizeTicket implements ShouldQueue
{
    use Queueable;

    public int $tries = 3;
    public int $timeout = 180;

    public function __construct(public int $ticketId) {}

    public function backoff(): array
    {
        return [10, 60, 300];
    }

    public function handle(ClaudeClient $claude): void
    {
        $ticket = Ticket::findOrFail($this->ticketId);

        $result = $claude->complete(
            system: view('prompts.ticket-summary-v3')->render(),
            user: $ticket->body,
            maxTokens: 1024,
        );

        $ticket->update(['ai_summary' => $result['text']]);
    }
}

Three settings matter. The job timeout must be longer than the slowest call you expect. The queue connection's retry_after must be longer than the job timeout, or a second worker will pick up a job that is still running. And retries should back off, because the most common failure is a rate limit, which more requests make worse. Run AI jobs on their own queue so a backlog of summaries never delays password reset emails.

How do you stream ChatGPT or Claude responses in Laravel?

Stream the provider's response to the browser over Server-Sent Events, so the first words appear in about a second instead of after the full answer. Ask the provider for a stream, read it line by line, and forward only the text deltas.

use Illuminate\Http\Request;
use Illuminate\Support\Facades\Http;

Route::post('/assistant/stream', function (Request $request) {
    $question = $request->validate(['q' => 'required|string|max:4000'])['q'];

    return response()->stream(function () use ($question) {
        $upstream = Http::withHeaders([
                'x-api-key' => config('services.anthropic.key'),
                'anthropic-version' => '2023-06-01',
            ])
            ->withOptions(['stream' => true])
            ->timeout(120)
            ->post('https://api.anthropic.com/v1/messages', [
                'model' => config('services.anthropic.model'),
                'max_tokens' => 4096,
                'stream' => true,
                'messages' => [['role' => 'user', 'content' => $question]],
            ]);

        $body = $upstream->toPsrResponse()->getBody();
        $buffer = '';

        while (! $body->eof()) {
            $buffer .= $body->read(1024);
            while (($pos = strpos($buffer, "\n")) !== false) {
                $line = trim(substr($buffer, 0, $pos));
                $buffer = substr($buffer, $pos + 1);
                if (! str_starts_with($line, 'data: ')) {
                    continue;
                }
                $event = json_decode(substr($line, 6), true);
                if (($event['type'] ?? '') === 'content_block_delta'
                    && ($event['delta']['type'] ?? '') === 'text_delta') {
                    echo 'data: '.json_encode(['text' => $event['delta']['text']])."\n\n";
                    if (ob_get_level() > 0) { ob_flush(); }
                    flush();
                }
            }
        }
        echo "event: done\ndata: {}\n\n";
        flush();
    }, 200, [
        'Content-Type' => 'text/event-stream',
        'Cache-Control' => 'no-cache',
        'X-Accel-Buffering' => 'no',
    ]);
})->middleware(['auth', 'throttle:ai']);

The X-Accel-Buffering header stops nginx from holding the stream back. Recent Laravel versions also include a helper for event streams; check the current docs before writing your own loop. OpenAI's streaming works the same way with different event names (response.output_text.delta). With PHP-FPM, each open stream holds a worker, so size your pool for concurrent chats or run streaming endpoints on Octane.

Where should you store prompts and AI call logs?

Store prompts as versioned templates and log every call to a database table with the prompt version, model, token counts, latency and cost. Without that log you cannot explain a bad answer, a cost spike or whether last week's prompt change helped.

Schema::create('ai_calls', function (Blueprint $table) {
    $table->id();
    $table->foreignId('user_id')->nullable()->index();
    $table->foreignId('tenant_id')->nullable()->index();
    $table->string('feature');
    $table->string('prompt_version');
    $table->string('provider');
    $table->string('model');
    $table->unsignedInteger('input_tokens')->default(0);
    $table->unsignedInteger('output_tokens')->default(0);
    $table->unsignedInteger('latency_ms')->default(0);
    $table->decimal('cost_usd', 10, 6)->default(0);
    $table->json('meta')->nullable();
    $table->timestamps();
});

Keep prompt templates in Blade views or a prompts table, named with a version (ticket-summary-v3), and never edit a version in place. Storing prompts in the database lets you tune tone without a deploy; we used that pattern in the AI Grief Companion, where prompts live in the database with a version history. Mask personal data before logging full inputs, and apply the same retention rules as the rest of your app.

How do you add RAG to Laravel with pgvector?

If your app runs on PostgreSQL, add the pgvector extension and keep embeddings in a normal table next to tenant IDs and permissions. Retrieval becomes one SQL query that filters by what the user may see and orders by vector similarity.

// migration
DB::statement('CREATE EXTENSION IF NOT EXISTS vector');
Schema::create('chunks', function (Blueprint $table) {
    $table->id();
    $table->foreignId('tenant_id')->index();
    $table->foreignId('document_id')->index();
    $table->text('content');
    $table->timestamps();
});
// dimension must match your embedding model
DB::statement('ALTER TABLE chunks ADD COLUMN embedding vector(1536)');
DB::statement('CREATE INDEX chunks_embedding_idx ON chunks USING hnsw (embedding vector_cosine_ops)');

Create embeddings with the OpenAI embeddings endpoint (Anthropic does not offer its own embedding models, so Claude projects use OpenAI or another embedding provider), then search:

$embedding = Http::withToken(config('services.openai.key'))
    ->post('https://api.openai.com/v1/embeddings', [
        'model' => config('services.openai.embedding_model'),
        'input' => $question,
    ])
    ->throw()
    ->json('data.0.embedding');

$chunks = DB::select(
    'SELECT id, document_id, content
       FROM chunks
      WHERE tenant_id = ?
      ORDER BY embedding <=> CAST(? AS vector)
      LIMIT 8',
    [$tenantId, '['.implode(',', $embedding).']']
);

Put the retrieved chunks into the prompt with their IDs and ask the model to cite them. The tenant filter in the WHERE clause is the important line: permissions are enforced before anything reaches the model. For chunking, hybrid search and evaluation, see our RAG implementation guide.

How do you control AI costs and rate limits in Laravel?

Use Laravel's rate limiter per user and per tenant, cache repeated results, and route simple steps to a smaller model. One looping script or one enthusiastic customer should never produce a surprise invoice.

// AppServiceProvider::boot()
RateLimiter::for('ai', function (Request $request) {
    return [
        Limit::perMinute(10)->by('user:'.$request->user()->id),
        Limit::perDay(2000)->by('tenant:'.$request->user()->tenant_id),
    ];
});

Other levers, in the order they usually pay off:

  • Prompt caching. Keep the long, stable part of the prompt (instructions, examples, reference text) first. OpenAI caches long repeated prefixes automatically; Claude caches what you mark with cache breakpoints. Cached input is billed at a large discount on both.
  • Model routing. A small model for classification and extraction, a large one only where it measurably helps.
  • Result caching. Cache::remember keyed by a hash of the input and prompt version for deterministic tasks.
  • Batch APIs for nightly jobs, at a discount on both providers.
  • Monthly budgets per tenant, checked against the ai_calls table before a call.

Typical running costs we see in 2026: tens of dollars a month for a low-traffic internal tool, up to several thousand for a busy customer-facing assistant.

How do you test OpenAI and Claude calls in Laravel?

Fake the HTTP layer so tests are fast, free and deterministic, and test what your code does with responses: parsing, saving, permission checks and error handling.

use Illuminate\Support\Facades\Http;

it('stores the ai summary on the ticket', function () {
    Http::preventStrayRequests();
    Http::fake([
        'api.anthropic.com/*' => Http::response([
            'content' => [['type' => 'text', 'text' => 'Customer cannot log in after reset.']],
            'usage' => ['input_tokens' => 120, 'output_tokens' => 11],
        ]),
    ]);

    $ticket = Ticket::factory()->create();
    (new SummarizeTicket($ticket->id))->handle(app(ClaudeClient::class));

    expect($ticket->fresh()->ai_summary)->toBe('Customer cannot log in after reset.');
    Http::assertSent(fn ($request) => $request['max_tokens'] === 1024);
});

Also test the failure paths: a 429 response, a timeout, an empty answer. Unit tests do not measure answer quality. For that, keep an evaluation set of 50 to 200 real inputs with expected outputs and run it on every prompt or model change, in CI or as an artisan command.

Which approach do we recommend?

For most Laravel apps, our verdict is: start with Laravel's HTTP client behind one service class, queue everything non-interactive, stream chat, log every call, use pgvector before any separate vector database, and add an SDK package only when you need features the raw API makes awkward. That stack is easy to test and easy to switch between providers.

We built our own AI SaaS, AI Resume Master, on PHP and the OpenAI API in about three months, and it reached 50,000 monthly active users. Symfony is our default framework for new products, but we work inside existing Laravel codebases without pushing a migration. If your team is on Symfony instead, read how to add AI to a Symfony app.

Next step

If you want a senior PHP team to add an OpenAI or Claude feature to your Laravel app, see our Laravel development services. Most first features take 3 to 8 weeks and start from $10,000, with a fixed quote after a short scoping call. Contact us with a description of the feature and the data it needs.

Case study

Domande frequenti

Both work in production. Laravel's HTTP client gives you full control, no extra dependency and easy testing with Http::fake, which is enough for one or two features. An SDK package adds typed responses, streaming helpers and provider-specific features. Anthropic publishes an official PHP SDK, and community packages cover OpenAI and multi-provider use. Whichever you pick, wrap it in your own service class so the rest of the app does not depend on it.

Move anything that does not need an instant answer into a queued job: summaries, classification, enrichment, document processing. Give the job a timeout above the slowest expected call, a small number of retries with backoff, and make sure the queue connection's retry_after is longer than the job timeout. For interactive chat, stream the answer so the first words appear within about a second.

Yes, if you run PostgreSQL. The pgvector extension stores embeddings in a normal table next to your tenant IDs and permissions, and an HNSW index keeps similarity search fast for up to several million chunks. You query it with plain SQL through Laravel's database layer. A dedicated vector database only becomes worth it at much larger scale or with special search needs.

Use Http::fake to return a recorded provider response, Http::preventStrayRequests so no test reaches the real API, and assertions on the request body to check the model, prompt version and limits you send. Test your parsing, error handling and permission logic this way. Quality of the model's answers is a separate concern, measured with an evaluation set of real examples, not with unit tests.

Two numbers matter: the build and the monthly API bill. With us, an AI feature in an existing Laravel app usually takes 3 to 8 weeks and starts from $10,000, with a fixed quote after a short scoping call. The API bill ranges from tens of dollars a month for a low-traffic internal tool to several thousand for a busy customer-facing assistant, and caching and model routing reduce it.

Avviamo il vostro progetto
Prenota una call