LLM features in production
RAG, structured output, streaming, and prompt orchestration — integrated into your React/Node product and built to survive real traffic.

I'm an ML & Applied AI Engineer and full-stack developer. I build production ML pipelines, integrate LLMs, develop agentic systems, and ship RAG pipelines into live products — alongside the search, auth, and infrastructure they run on. I work primarily as a fractional / part-time Applied AI engineer (~10–15h/week) for EU and US teams that want production AI features shipped steadily, without hiring a full-time headcount.
LLM integration, agentic systems, and RAG in React/Node. Full-stack is how I deliver — not a separate service line.
RAG, structured output, streaming, and prompt orchestration — integrated into your React/Node product and built to survive real traffic.
AI agents with tool use, function calling, and autonomous workflows that take real actions inside your product — not just generate text.
React, Next.js, Node.js, vector DBs, deployment — I own the stack end to end so your AI features actually ship, not stall in a notebook.
From retrieval and structured LLM output to agents that call tools and automate workflows — engineered for production in the JavaScript stack.
Ground LLMs in your data with embeddings, vector search, and retrieval tuned for your domain.
Structured output, streaming, prompt orchestration, and safe integrations that survive production traffic.
Agents with tool use, function calling, and multi-step reasoning that automate real workflows inside your codebase and product.
Testing, evaluation, and monitoring so AI features stay accurate after launch.
Real products with real users — numbers included, not just adjectives. AI-led first: najdiavto.com ships a natural-language search feature on top of a production marketplace.
Production car marketplace with real dealer inventory — faceted search plus natural-language AI search across fuel, transmission, body type, color, doors, and Euro norm.
View case studyHosted auth infrastructure for mobile and web — OAuth, short-lived JWTs with rotating refresh tokens, hashed sessions, role-based user management.
Visit siteBrowser PDF workspace — upload, edit, annotate, and sign documents with no install, no ads, and a proper landing page.
Visit siteOpen-source ML projects that demonstrate specific techniques — from NLP to recommendation systems.
End-to-end ML service for a car marketplace — price intelligence, hybrid search, learning to rank, and MLOps with monitoring and safe retraining.
View case studyMultilingual NLP demo — sentiment analysis and text classification in English, Serbian, and German with a feedback loop for continuous improvement.
View case studyHybrid movie recommendation engine — collaborative filtering (SVD) + content-based genre similarity. Powered by MovieLens with 9,700+ films.
View case studyAutomated ML pipeline — upload data, train regression models (LinearRegression, Ridge, RandomForest), compare metrics, and make predictions. No code needed.
View case studyYou work with me directly — no sales team, no account manager. I scope honestly, ship weekly, and care about what happens after launch.
A focused kickoff: review your data and product, pin down the highest-leverage AI feature, and agree on metrics that define success.
Within the first cycle I deliver something real in production — a RAG endpoint, a structured LLM feature, or an automated workflow — not a slide deck.
We decide together: a steady fractional retainer (~10–15h/week) to keep shipping, or a scoped fixed project. No long-term lock-in.
Tell me what you're building — I'll get back within 24 hours. No obligations, fully confidential. Open to fractional / part-time retainers (~10–15h/week) across the EU and US.