SpaceWhale is a global IT company. We develop and provide tech solutions that change everyday lives, transform industry insight, and create leading-edge products. As we continue to grow, we’re building a diverse team where every member can explore, learn, and flourish. We strive for perfection while putting our people first.
18 травня 2026

QA Engineer (AI) (вакансія неактивна)

Київ

About the Role:
We are hiring the foundational QA for XBO’s AI department. You will own quality across products built by AI agents — both standalone company products and modules across our various platforms — by combining traditional manual and automated testing with AI-driven testing approaches you design and build yourself. This is not a classical QA role: it is a hands-on lead position with full ownership of how quality is defined, measured, and shipped.

How We Work:

  • Day Agent and Night Agent: every team member (QA included) runs two primary agents — a Day Agent that works alongside us and a Night Agent that runs in the background. Both can spawn sub-agents.
  • Two daily iterations: a morning sync where each member outlines the day’s tasks, assigns them to their Day Agent, and reviews what the Night Agent produced overnight. Reviewed work then moves into the next assignment cycle.
  • Planner and reviewer, not writer: we do not write code manually by default. We plan, delegate, review, refine, and continuously train the agents to be smarter on each iteration.

Responsibilities:

  • Own the testing strategy for every product and module — deciding what gets tested manually, what gets automated, and what is delegated to AI agents.
  • Run exploratory, regression, and UAT testing on critical flows (financial flows, security, third-party integrations) before each release.
  • Build and maintain end-to-end automation suites (Playwright / Cypress / Selenium), API testing (REST and GraphQL), and database validation.
  • Build a Day testing agent and a Night testing agent that run suites, analyze results, and open tickets autonomously.
  • Run specialized AI testing: hallucination detection, prompt injection, RAG retrieval quality, multi-agent failure modes, jailbreak resistance.
  • Define and operate evals and benchmarks for models and agents; track quality over time.
  • Set quality gates for production deployment and own end-to-end bug lifecycle: identification, reproduction, root cause, verification.
  • Work directly with internal engineers, external development partners, product, and business stakeholders; report on quality posture continuously.

Required:

  • Experience: minimum 4 years in QA with at least 2 years of end-to-end ownership (not just executing test cases someone else wrote).
  • Manual + automation: strong fundamentals in exploratory, regression, and UAT; hands-on with at least one of Playwright, Cypress, or Selenium; code-based test suites, not record-and-playback.
  • API + DB: REST/GraphQL testing, Postman, OpenAPI/Swagger; strong SQL with PostgreSQL or MySQL.
  • CI/CD: comfortable adding test stages to GitHub Actions / Jenkins / GitLab pipelines.
  • Claude Code: hands-on experience, or a clear demonstrated ability to ramp up to daily working proficiency within 2–4 weeks. “I played with it” is not enough.
  • LLMs and prompt engineering: advanced-user-level understanding of context windows, prompting, hallucinations, and AI failure modes; ability to write structured prompts and improve them systematically.
  • Python at scripting level: able to write a script that calls an API, processes the response, and validates an outcome (requests, pytest, basic data structures).
  • Languages: Hebrew at native or near-native level; professional English for working with external partners and technical documentation.

Nice to Have:

  • Demonstrated experience testing LLM / GenAI / chatbot products in production.
  • Familiarity with RAG, vector databases (ChromaDB and similar), and multi-agent systems.
  • Security testing background (OWASP Top 10, dependency scanning, secret scanning).
  • Familiarity with MCP (Model Context Protocol) and agent frameworks.
  • Performance testing (k6, JMeter, Locust); experience in regulated environments.

Who We’re Looking For:
This is a foundational hire with no embedded mentoring.
We need someone who:

  • Takes loosely defined problems, breaks them down, and executes without being walked through every step.
  • Owns quality end-to-end — “I am responsible for what ships”, not “I test what I was asked to test.”
  • Learns fast and independently. Tools change weekly; people who cannot keep up will not last.
  • Communicates clearly with technical and non-technical stakeholders, and can explain risk and trade-offs in plain language.