The challenge behind the engagement.
This service is for teams shipping AI features through an automated delivery pipeline who want security to run as pipeline steps, not a final review. We baseline your current lifecycle against recognized secure AI development practices and attach a design-time requirements checklist at intake: input and output handling, tool permissions, and logging. We threat model the flows that carry prompts, retrieved content, tool calls, and memory, and turn the results into acceptance criteria. We then specify the gates: scanning model files for safe formats, verifying their signatures, generating a bill of materials and model cards, secret scanning, injection and leakage tests with pass thresholds, and a manual review gate for any new tool or wider permission. Each gate writes an artifact your auditors can use.
For engineering managers, AppSec leads, and platform teams with two or more AI features in flight and a release cadence they do not want to slow. Engage when the first agent tool is being added, when fine-tuning starts in-house, or when SOC 2, ISO/IEC 42001, or CMMC evidence for AI changes is due.
What we do.
What you can use.
Baseline the pipeline you already run
We start with read access to your repositories and pipelines, two or three sample AI features with their prompts, your current threat modeling and code review practice, the model and provider list, and any test datasets you hold. We map what you do today against recognized practices for model provenance, training-data protection, pre-release evaluation, and secure deployment, and rank the gaps.
Model threats and write the gates
We complete two or three threat models on your own features and leave the template behind. Then we write the gate specification for your pipeline. Automated gates cover unsafe-format detection, safe-format enforcement, model signature checks, a bill of materials and model cards, prompt and policy linting, injection and leakage tests with thresholds, and guardrail drift checks. A manual review step gates any new tool.
Wire gates, train engineers, map evidence
Your engineers wire the gates into the first pipeline with us; we review the run logs and adjust thresholds. You receive the gap assessment and backlog, requirements checklist, threat model template with completed models, gate specification, and eval and red-team test plan. An artifact map links each gate to SOC 2, ISO/IEC 42001, or CMMC evidence, and we train your engineers.
Who it’s for.
When you need it.
- Engineering and AppSec teams shipping AI features on a regular release cadence
- Platform groups standardizing how every team ships AI changes safely
- Organizations fine-tuning or training models in house rather than only calling APIs
- Teams that must produce audit evidence for AI changes on each release
- The first agent tool or write-capable integration is being added to a feature
- In-house fine-tuning or model training is starting for the first time
- SOC 2, ISO/IEC 42001, or CMMC evidence for AI changes is due
- Security reviews keep arriving late, after features are already built
What the scope can include.
- 01
Gap assessment of the current lifecycle against recognized secure AI development and AI risk-management practices
- 02
Design-time AI security requirements attached to intake, with misuse and abuse cases written per feature
- 03
Threat models on prompt, retrieval, tool, and memory flows, mapped to recognized AI and agent risk catalogs, converted to acceptance criteria
- 04
Pipeline gates: model-file scanning, signature verification, bill-of-materials and model card generation, secret scanning, and prompt and policy linting
- 05
Eval suites for prompt injection and data leakage with thresholds, guardrail drift checks, and a manual review gate for new tools or widened permissions
- 06
Runtime feedback loop: incident and test findings written back as regression tests, aligned with your incident-response plan
What it typically costs.
One rate: $150/hour.
Every engagement is priced by the hours it takes at one flat rate, with scoping, the work, and the final deliverables included. Find the size closest to yours.
One development team shipping one or two AI features through one release pipeline
About 48–72 hoursA few teams with several AI features, three threat models, and checks built into the first pipeline
About 96–140 hoursCompany-wide rollout across many pipelines, including in-house model training and audit evidence
About 160–240 hours- Number of teams, pipelines, and AI features in flight
- Whether you train or fine-tune your own models
- How many threat models and release checks are built with your engineers
- Audit frameworks the evidence must map to (SOC 2, ISO 42001, CMMC)
Ranges are planning estimates at $150/hour, not a quote. Your price is confirmed in writing after a scoping call, before any work begins.
What you take forward.
- Gap assessment against recognized secure AI development practices, with a prioritized backlog and named owners
- AI security requirements checklist, threat model template, and two or three completed threat models on your features
- Pipeline gate specification with reference configuration for your own delivery pipeline, plus a test and red-team plan with pass thresholds
- Artifact and evidence map linking each gate to the SOC 2, ISO/IEC 42001, or CMMC evidence it produces
Final coverage, deliverables, timing, and any retesting or implementation work are confirmed before the engagement begins.
Before we get started.
How do we test prompt injection in CI when model output is nondeterministic?
By counting rather than asserting. An injection test runs each payload several times against the deployed prompt and tools and records how many attempts succeed; the gate fails when the success rate crosses a threshold you set per feature, not on a single response. We pin the model version and temperature for the test, keep the payload set in the repository so it grows with every incident, and run the suite in parallel with unit tests so it rarely extends the build. Test scores are a regression signal, not proof of safety; the manual gate for new tools and the AI Security Testing core cover what tests cannot.
Is penetration testing included, or is that a separate engagement?
Separate. This service builds the requirements, threat models, and gates that run on every change, and it hands over a test and red-team plan your engineers execute in the pipeline. Manual, attacker-driven testing of a finished feature, where the principal attempts injection, tool abuse, and retrieval leaks by hand and reports reproduction transcripts, is the AI Security Testing core. Many clients sequence them: lifecycle integration first so findings from a later test become regression tests, then an LLM Application Penetration Test or Agentic AI Security Assessment before launch. Gates do not replace that test; they keep its findings from coming back.
- Secure Software Development Practices for Generative AI and Dual-Use Foundation Models: An SSDF Community Profile (NIST SP 800-218A)
- AI Risk Management Framework
- Joint Guidance on Deploying AI Systems Securely
- AI Data Security: Best Practices for Securing Data Used to Train & Operate AI Systems
- OWASP Top 10 for Agentic Applications for 2026
- Model Context Protocol Specification: Authorization
- Protecting Controlled Unclassified Information in Nonfederal Systems and Organizations (NIST SP 800-171 Rev. 3)
