AI Evaluation Engineer

siteground· AI Products
Apply Now ↗
📍 SofiaFull time

About this role

YOUR ROLE:

    At SiteGround, we’ve built and launched our own suite of AI products:

    Coderick AI — our vibe coding platform.
    AI Studio — access to flagship AI models, providers, and intelligent AI Agents.
    WordPress AI Agent — one AI assistant for virtually every WordPress task.

    We’re moving fast: adding new capabilities, launching new AI agents, and serving a growing number of customers who use our AI products every day.

    This role sits at the heart of how we measure and improve those products. Your job is to make our LLM-powered features provably reliable — not just impressive in a demo.

    Prompt engineering is a central part of the role. You’ll design, test, and iterate on prompts and agent behaviour directly. But your defining contribution will be rigour: turning “this seems to work” into measurable, repeatable, production-ready behaviour through evaluation and observability.

    We work across the major model ecosystems — OpenAI, Google Gemini, and Anthropic — and you’ll regularly make pragmatic decisions about model choice, quality, latency, and cost.

    You’ll work closely with backend engineers, fellow evaluation engineers, and product managers. Together, you’ll dig into real customer interactions, identify where our AI succeeds or fails, and turn those insights into better prompts, agents, tools, and product experiences.

YOUR RESPONSIBILITIES:

    Evaluation & Observability

  • Build high-quality evaluation datasets with ground truth and define task-specific metrics such as correctness, task completion, efficiency, cost, token usage, and safety;
  • Run model bake-offs and before/after prompt experiments — every meaningful prompt change should be a measured experiment, not a guess;
  • Design evaluations for harder cases where there is no clean ground truth, including agentic workflows, tool-calling correctness, and structured-output validity;
  • Own production observability through trace analysis, per-tool dashboards, error classification, and session-level review;
  • Root-cause failures, identify recurring patterns, understand their impact on quality and cost, and turn findings into prompt improvements and engineering tickets;
  • Help close the loop between users and product teams by translating real customer behaviour and product questions into clear, measurable insights about how our AI features are used and where they can improve.

  • Prompting & Agent Engineering

  • Design, version, test, and continuously improve system prompts for domain-specific agents;
  • Architect agent behaviour, including tool orchestration, multi-turn flows, and safeguards such as plan-before-execute, read-before-write, and always-confirm for destructive actions;
  • Build reusable, modular prompt components for areas such as safety, tone, context, and tool documentation;
  • Select the right model for each task, balancing quality, latency, and cost.

  • Tooling & Integrations

  • Work across the API layer our AI agents act through, integrating the REST endpoints that expose application data and operations to the LLM;
  • Help design and implement guardrails that prevent agents from corrupting user data, including permission enforcement and safe defaults;
  • Debug the full request path end to end — from an agent’s decision and tool call to the live result — correlating application logs with agent traces.

OUR EXPECTATIONS:

  • Strong prompt engineering and LLM application experience — you’ve shipped agents or LLM-powered features to production, not just prototyped them;
  • Hands-on evaluation experience and a data-driven mindset — you’ve built evaluation datasets and metrics, and you measure results before claiming something works;
  • Python experience and comfort consuming and integrating REST APIs;
  • Familiarity with modern LLM tooling and patterns, including function/tool calling, RAG, agentic workflows, observability platforms, and current model families;
  • Healthy scepticism toward AI output and a strong instinct for safety, failure modes, and guardrail design.

GREAT ADVANTAGE WILL BE:

  • Experiment design and statistics literacy, including A/B testing, significance, and sampling, so your evaluation conclusions hold up under scrutiny;
  • Experience designing LLM-as-a-judge evaluations and calibrating automated judges against human labels;
  • Red-teaming and adversarial testing experience, including prompt injection, jailbreaks, and structured-output abuse;
  • Comfort working with data tools for exploring and slicing evaluation results, such as SQL, pandas, DuckDB, notebooks, or similar;
  • Experience with platforms/solutions such as Langfuse, Langchain, n8n, Google ADK.

WHAT WE OFFER:

    At SiteGround, we work hard and challenge ourselves to exceed expectations. To support your drive, we provide a thoughtfully designed benefits package created to help you thrive both at work and beyond.

     

  • Competitive remuneration and a performance-based bonus system;
  • Shorter workday on Fridays – we finish at 3:00 PM;
  • 5 extra days off at 5 years of service, plus 1 more each year after – up to 10;
  • 2 extra paid days off for volunteering to support the causes you care about;
  • Premium additional health insurance coverage and annual medical check-ups;
  • Modern and cozy offices across Bulgaria, plus hybrid work options;
  • Free parking and a metro shuttle;
  • Chef-prepared breakfast and lunch provided daily at our HQ restaurant – all on us;
  • Fully equipped in-office gym with expert fitness guidance;
  • On-site sports at our HQ: yoga, spinning, Brazilian Jiu-Jitsu, dancing, table tennis, and more;
  • Multisport or Coolfit cards fully covered by the company;
  • Free massages;
  • Memorable team events and internal gatherings;
  • Knowledge-sharing culture through meetups, events and conferences;
  • Free employee access to the SiteGround products;
  • Anniversary and company gifts along the way.

Come and make a difference at SiteGround with people who inspire you to grow!

Please note: Only shortlisted candidates will be contacted for further steps.

 

Frequently Asked Questions

Is the salary disclosed for the AI Evaluation Engineer position at siteground?
The salary for this AI Evaluation Engineer role at siteground is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the AI Evaluation Engineer position at siteground located?
This AI Evaluation Engineer role at siteground is based in Sofia. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the AI Evaluation Engineer role at siteground full-time or part-time?
This is listed as a Full time position. It is posted as a AI Evaluation Engineer role in the AI Products department at siteground.
Which team or department does the AI Evaluation Engineer at siteground belong to?
This AI Evaluation Engineer position is part of the AI Products department at siteground. See the full job description for more information about the team structure and responsibilities.
How do I apply for the AI Evaluation Engineer position at siteground?
Click the "Apply Now" button on this page. You will be redirected to siteground's official application portal hosted on lever where you can submit your application directly.
When was the AI Evaluation Engineer job at siteground posted?
This AI Evaluation Engineer position at siteground was posted on Sep 21, 2026. Apply as soon as possible — early applications are often reviewed first.
AI Evaluation Engineer
siteground
Apply for this role ↗

You'll be redirected to siteground's official application page on Lever.