Software Quality Assurance Engineer – GenAI / LLM
Job Description
Primary Focus: Software Quality Assurance / Software Testing / Test Automation – GenAI, LLM & Agentic AI Secondary Exposure: Solution Analysis / Technology Solution Design / Enterprise Integration Domain / Project: Global Markets, Capital Markets Banking Technology & Market Risk Technology
Role Overview We are looking for a Senior GenAI Quality Engineer / Solution Analyst to design, analyse, test and validate production-grade Generative AI (GenAI), Large Language Model (LLM), RAG and Agentic AI applications within a complex enterprise environment. This is not a traditional manual QA or software testing role. The role combines:
- Software Quality Engineering
- GenAI / LLM Testing & Evaluation
- Agentic AI / AI Agent Testing
- UI & API Testing
- Test Automation
- Solution Analysis
- Enterprise Integration Testing
- Observability & Troubleshooting
You will work across discovery, solution design, development, testing and release, translating business requirements into clear application behaviours and validating end-to-end application quality across user interfaces, APIs, data flows, LLMs, RAG components, AI agents and enterprise integrations. Key Responsibilities GenAI / LLM Quality Engineering
- Define and execute end-to-end quality engineering and test strategies covering:
- Web / UI workflows
- REST APIs
- Backend services
- Enterprise integrations
- GenAI applications
- LLM workflows
- RAG pipelines
- Agentic AI / AI Agent interfaces
- Perform GenAI / LLM testing and evaluation covering:
- Response quality
- Task completion
- Grounding
- Faithfulness
- Relevance
- Consistency
- Citation accuracy
- Hallucination risk
- Safe failure behaviour
- Test non-deterministic / probabilistic AI systems using:
- Evaluation datasets
- Repeat testing
- Quality thresholds
- Acceptance criteria
- Regression evaluation
- Validate RAG / Retrieval-Augmented Generation solutions, including retrieval quality, grounding and response accuracy.
Agentic AI / AI Agent Testing Test end-to-end Agentic AI and AI Agent workflows, including:
- Multi-turn conversations
- Context handling
- Agent planning
- Tool selection
- Tool calling / function calling
- Tool inputs and outputs
- State transitions
- Memory and state
- Human-in-the-loop approvals
- Handoffs
- Retries
- Timeouts
- Fallback behaviour
- Error recovery
- Termination conditions
- Partial failures
Validate that AI agents behave correctly across both successful and failure scenarios. Software & API Quality Engineering Perform:
- Functional Testing
- Integration Testing
- API Testing
- Regression Testing
- Exploratory Testing
- Negative Testing
- Resilience Testing
- Basic Performance Testing
- End-to-End Testing
Design comprehensive REST API tests covering:
- API contracts
- Authentication
- Authorisation
- Input validation
- Error handling
- Idempotency
- Rate limits
- Downstream system failures
Test web application behaviour across browsers and realistic end-user journeys, including:
- Loading states
- Interrupted sessions
- Error messages
- Feedback capture
- Accessibility fundamentals
Test Automation Develop and maintain risk-based test automation that reduces:
- Regression testing time
- Manual testing effort
- Release cycle time
- Production risk
Use automation frameworks and tools such as:
- Playwright
- Cypress
- Selenium
- pytest
- REST Assured
- Postman
- Equivalent UI / API automation frameworks
Apply pragmatic automation principles by prioritising stable, high-value and frequently executed test scenarios. GenAI Evaluation & AI Safety Testing Validate LLM and GenAI applications for:
- Grounded responses
- Hallucinations
- Retrieval quality
- Citation accuracy
- Prompt behaviour
- Prompt injection
- Unsupported requests
- Restricted content handling
- Safe failure behaviour
- Adversarial scenarios
Support AI evaluation / LLM evaluation using appropriate evaluation datasets, quality metrics and repeatable evaluation approaches. Exposure to AI Red Teaming / Adversarial Testing would be advantageous. Observability & Troubleshooting Use application and GenAI observability to identify the source of defects across:
- Application
- LLM / Model
- RAG / Retrieval
- Data
- API / Integration
- Platform
Analyse:
- Logs
- Distributed traces
- API requests / responses
- Payloads
- Network calls
- Database records
- Agent execution traces
Exposure to observability and LLM evaluation tools such as:
- Langfuse
- LangSmith
- OpenTelemetry
- Elastic / Elasticsearch
- Splunk
is advantageous. Solution Analysis & Design The role also acts as a hands-on Solution Analyst for GenAI applications. Responsibilities include:
- Partner with product owners, business users, architects, engineers and GenAI specialists during discovery and solution design.
- Analyse proposed GenAI use cases and determine whether the requirement should use:
- Conventional application logic
- Deterministic business rules
- Search / retrieval
- RAG
- Workflow automation
- Agentic AI
- Human approval
- Translate business requirements into:
- Functional requirements
- End-to-end solution flows
- User journeys
- Acceptance criteria
- Interface behaviour
- Decision rules
- Non-functional requirements
- Map interactions across:
- User Interfaces
- APIs
- LLMs / Models
- Prompts
- RAG / Retrieval components
- Enterprise data sources
- AI Agent tools
- Downstream enterprise systems
- Analyse solution design trade-offs involving:
- Quality
- Complexity
- Cost
- Latency
- Security
- Data access
- Maintainability
- Operational risk
- Identify missing controls, integration assumptions, ownership gaps, failure scenarios and operational risks before development begins.
- Support the design of:
- Human-in-the-loop approval
- Fallback flows
- Escalation
- Exception handling
Solution Documentation Produce practical technical and functional artefacts including:
- Process Flows
- Sequence Diagrams
- Context Diagrams
- Interface Specifications
- Decision Tables
- User Stories
- Acceptance Criteria
- Test Scenarios
- Traceability Documentation
Maintain traceability across: Business Requirement → Solution Design → Implementation → Test / Evaluation Scenario → Release Evidence Release Quality & Governance Create and maintain:
- Test scenarios
- Test datasets
- Reusable regression scenarios
- Test evidence
- Defect reports
- Quality metrics
- Release quality reports
Provide evidence-based release recommendations identifying:
- Known defects
- Known limitations
- Residual risks
- Quality concerns
- Areas requiring production monitoring
Core Requirements Experience
- 5–8 years of experience in Software Quality Engineering, Test Engineering, Test Automation, SDET or similar hands-on software testing roles.
- Strong experience testing complex enterprise applications.
- Strong experience testing:
- Web applications
- REST APIs
- Backend services
- Enterprise integrations
Test Automation / Programming Hands-on experience with one or more of:
- Playwright
- Cypress
- Selenium
- pytest
- REST Assured
- Postman
- Equivalent automation frameworks
Working programming knowledge of:
- Python
- Java
- JavaScript
- TypeScript
Candidates should be capable of developing, reviewing and troubleshooting test automation. Software Engineering / DevOps Experience with:
- Git
- Pull Requests
- CI/CD
- Automated Testing
- Test Reporting
- Defect Management
Experience validating distributed systems including:
- Asynchronous Processing
- Queues
- Batch Processing
- APIs
- Downstream Dependencies
- Enterprise Integrations
GenAI / LLM Requirements Practical understanding of:
- Generative AI / GenAI
- Large Language Models / LLM
- LLM Evaluation
- LLM Testing
- Retrieval-Augmented Generation / RAG
- RAG Evaluation
- Agentic AI
- AI Agents
- Multi-Agent Workflows
- Prompts / Prompt Engineering
- Context Windows
- Embeddings
- Tool Calling
- Agent Memory & State
- LLM Observability
Candidates should understand how GenAI applications differ from conventional deterministic software and how to validate probabilistic AI behaviour. Security & Risk Testing Understanding of software and GenAI security fundamentals including:
- Access Control
- Authentication / Authorisation
- Sensitive Data Handling
- Input Validation
- Auditability
- Prompt Injection
- AI Safety Testing
- Adversarial Testing
Nice to Have Experience with:
- Banking / Financial Services
- Regulated enterprise environments
- Contract Testing
- Service Virtualisation
- Synthetic Monitoring
- Performance Testing
- AI Red Teaming
- Accessibility Testing / WCAG
- Kubernetes
- OpenShift
- AWS
- Containerised Application Deployment
Key Domain / Technical Skills
- Software Quality Engineering, API Testing & Test Automation
- GenAI / LLM Evaluation, RAG & Agentic AI Testing
- Solution Analysis, Observability & Enterprise Integration Key Search Keywords GenAI Quality Engineer,AI Quality Engineer,LLM Quality Engineer,Generative AI Testing,GenAI Testing,LLM Testing,LLM Evaluation,AI Evaluation,Agentic AI Testing,AI Agent Testing,RAG Testing,RAG Evaluation,Retrieval-Augmented Generation,Software Quality Engineering,Quality Engineering,Software QA,Test Automation,SDET,Automation Testing,API Testing,REST API Testing,UI Testing,Integration Testing,Regression Testing,End-to-End Testing,Playwright,Cypress,Selenium,pytest,REST Assured,Postman,Python,Java,JavaScript,TypeScript,CI/CD,Git,Prompt Testing,Prompt Injection,Hallucination Testing,Grounding,Faithfulness,AI Safety Testing,Adversarial Testing,AI Red Teaming,Langfuse,LangSmith,OpenTelemetry,Elastic,Splunk,Observability,Distributed Systems,Kubernetes,OpenShift,Solution Analysis