---
title: "A Deploy Job That Fetches Its Own Test Data | Clarive"
description: "Many teams still test against a copy of production that someone restored into staging over a weekend. We move that work into the deploy job, where a step in a Clarive pipeline rule calls a Gigantics pipeline that masks and trims production data and loads it into the environment about to be tested."
url: "https://clarive.com/blog/where-the-test-data-comes-from"
published: "2026-09-30"
author: "Rodrigo Gonzalez"
---

# A Deploy Job That Fetches Its Own Test Data

A Clarive pipeline rule can call a Gigantics pipeline before each release, so the environment being tested gets a masked cut of production data taken at deploy time instead of an old restore.

Software teams have spent more than a decade automating the way code gets to production. The data that code is tested against mostly still gets there by hand.

At a lot of companies it arrives as a restore of the production database into staging on a Sunday night. The masking is done by a script somebody wrote years ago, and the list of who can log in to staging lives in a spreadsheet. None of this runs inside the deployment pipeline. So none of it is reviewed or approved, nothing is logged and there is no record of which release the data was meant for.

We think the deploy job should produce the test data itself, in the same run that ships the code. In practice that is a few steps in a [Clarive pipeline rule](https://docs.clarive.com/en/hcl/rules/). Before the release goes out, the job calls a [Gigantics](https://www.gigantics.io/) pipeline. Gigantics reads the production database, anonymizes it, cuts it down to a smaller subset and loads the result into the database of whichever environment the job is deploying to.

The integration is a URL kept on a deploy server and one Clarive variable, plus a few blocks of Clarive HCL. Tests then run on data taken from production during that same deploy. Customer records reach the environment only after they have been masked.

Code has a fairly strict route to production. Every change is a commit and every build can be reproduced. A deployment needs an approval, and it leaves a log. Test data usually gets none of that, and most of the teams we talk to rely on one of three setups that each go wrong in a different way.

The most common is the weekend copy. A scheduled job restores production into staging, and by Monday the testers have realistic data to work with. They also have every customer the company holds, on servers that were never hardened the way production was. Each restore needs someone to sign off on it, so restores get rarer from one quarter to the next, and staging drifts further from production until a whole class of bugs stops reproducing there at all.

Hand-written fixtures are the safe choice. They are small and completely understood. They also have none of production's oddities: no customer with 400 orders, no street name with an apostrophe in it, no account that was migrated twice and still carries two IDs. The odd rows are usually where the bugs are.

The third option is a masking script run over the copy. It replaces names and card numbers in place and works fine until the day it rewrites a primary key in the `customers` table and misses the foreign key pointing at it from `orders`, at which point the test suite fails for reasons that have nothing to do with the release.

We took three requirements from those failures. The data has to be as current as production. Real customer records should exist only as long as an environment needs them. And the links between tables have to survive the masking.

A migration shows why the first one matters. Say a release adds a `NOT NULL` constraint to `orders.shipping_region`. On a copy restored three months ago, every order has a region, so the migration passes. In production, some of the orders taken since then don't have one, and the migration fails in the middle of the release window.

In our setup, provisioning runs at the start of the deploy. The dataset is cut from production as it stands when the deploy runs, and the release's own [database migrations](https://clarive.com/blog/database-release-automation) run on top of it in the same order they will run in production. So the migration gets tested too. A weekend copy can't do that. It shows production as it was on the day of the restore and gets a little more out of date every day after.

Exposure is the second problem. A standing copy of production is exposed for as long as it exists, which in practice tends to mean indefinitely. It sits on a server whose access list was last reviewed when the server was built. If an auditor asks who could see [customer data](https://clarive.com/blog/implementing-gdpr-made-more-simple) in staging in June, the honest answer is whoever had a login, and nobody can produce that list.

A dataset built by the deploy job has no such history. It is masked before it arrives, and it is replaced or dropped along with the environment it was made for. Security teams will know the idea from short-lived credentials. Here it is applied to rows instead of tokens.

![Timeline over four weeks: a continuous coral bar for the Sunday copy of production, and short separate indigo segments for datasets provisioned per deploy job.](https://clarive.com/images/blog/where-the-test-data-comes-from/exposure-window.webp)

A standing copy of production holds real customer data the whole time. A dataset provisioned by the deploy job exists only while its environment does.

Size is the third. Most test suites do not need four terabytes. They need a few thousand customers with the right kinds of cases among them, and anything past that costs disk space and restore time while engineers wait on both. So the data is cut down to a subset, and that is the step where simple tools break. Take 10,000 rows from `customers` and 10,000 from `orders` without checking which orders belong to which customers, and many of the orders will point at customers who are not there. When the tests fail, people blame the release.

## How the step is set up

The design splits the work between the two products. [Clarive](https://clarive.com/features) decides when data is needed and for which environment. Gigantics decides what the data looks like. The deploy job never learns how a column is masked, and Gigantics never learns what a release is. The only thing connecting them is a single URL.

Three pieces have to be configured. The first is a Gigantics rule, which defines the dataset. It says which tables and rows go in and which fields get anonymized or synthesized, and it sets whether the masking is deterministic.

Each environment then gets its own Gigantics pipeline, which runs the rule and loads the result into that environment's database. A pipeline can be started from a URL that carries an API key, and that URL is all the deploy job needs.

The last piece is a few steps in a Clarive pipeline rule. They call the pipeline for the job's target environment and check that the data arrived. Only after that does the release deploy.

![Diagram: a Clarive deploy job with build, provision data, deploy, test and promote stages. The provision data stage calls a Gigantics pipeline that runs a rule over the production tap using a dictionary, producing a versioned dataset that is loaded into DEV, QA or UAT databases.](https://clarive.com/images/blog/where-the-test-data-comes-from/deploy-job-architecture.webp)

The deploy job calls a Gigantics pipeline, which loads a masked dataset into the environment the job is targeting.

Work in Gigantics starts with discovery. Gigantics connects to a database, which it calls a tap, reads the schema and labels every field it believes holds personal data, such as names, email addresses, phone numbers and national ID numbers. People then confirm or correct the labels, and every change has to come with a reason. The changes go into an audit report that can be signed. Once it is signed, it can't be edited or deleted, so it records the state of the tap on that day.

Gigantics is installed on the customer's own servers. Its documentation asks for MongoDB 4.0 or later and recommends 16 gigabytes of RAM, and because of where it runs, the tap reads production from inside the company's own network.

A rule is a small program run against the tap. Its query operations pick the subset, using Include/Exclude, Where and Limit. Its data operations change the values. Anonymize swaps them for realistic fakes and Synthesize generates rows from the schema, while Transform runs JavaScript functions the user writes. Two of a rule's options matter most for testing. Determinism can be set to Random, or to Deterministic with a seed. Check foreign keys makes Gigantics track primary and foreign keys across the whole database, so the subset keeps its relationships.

Nearly every order-processing test suite depends on the rule that every order belongs to a customer who exists. A plain Limit can break that. Check foreign keys keeps it intact, though not for free. Gigantics caches data while it searches for master tables, and on a complicated schema the option can make the job noticeably slower. We leave it on anyway. A dataset that shows up quickly with orphaned orders in it costs more afternoons than it saves.

Deterministic masking turns each input into the same output every time. Customer 88213 becomes 40117 in `customers`, and in `orders` and `payments` as well. Gigantics keeps those mappings in dictionaries, and one dictionary can be shared by every tap and sink in a project. That keeps keys consistent across separate databases, which matters in older systems that store customers in one database and payments in another.

![Diagram: customer 88213, Marta Ruiz, in the customers, orders and payments tables of production passes through a dictionary and becomes customer 40117, Ana Soler, in all three tables of the delivered dataset.](https://clarive.com/images/blog/where-the-test-data-comes-from/referential-integrity.webp)

One shared dictionary maps customer 88213 to 40117 in every table, so the tables still join.

On the Clarive side, the steps go in the application's pipeline rule, written in [Clarive HCL](https://docs.clarive.com/en/hcl/overview/) and placed just ahead of the steps that ship the release. They run on a deploy server that can reach both Gigantics and the test databases. Clarive reaches it the way it reaches any other server, over SSH or through a Clarive agent, and in HCL it is a `generic_server`:

```hcl
generic_server "data-runner" {
  description = "Calls Gigantics and checks the test data"
  hostname    = "data-runner.example.invalid"
}
```

Each environment's pipeline URL sits in a file on that server, readable only by the account Clarive logs in as. A project variable tells the rule which file belongs to which environment. These are the lines this setup adds to the project:

```hcl
project "shop" {
  variables = {
    DEV = {
      gigantics_url_file = "/etc/gigantics/shop-dev.url"
    }
    QA  = {
      gigantics_url_file = "/etc/gigantics/shop-qa.url"
    }
    UAT = {
      gigantics_url_file = "/etc/gigantics/shop-uat.url"
    }
  }
}
```

There is no `gigantics_url_file` for PROD on purpose. This rule deploys to test environments only, and production goes through a rule without these steps.

That is not a reason to leave PROD out of the project's file, though. An import saves the project's variables whole, the way the Variables tab in the interface does, so [an environment left out of the file](https://docs.clarive.com/en/hcl/resources/#variables) is removed from the project. The real file has to keep every environment the project deploys to, PROD and the shared `"*"` values included, with these lines added to the ones already there.

The version of the rule we recommend looks like this, with the steps that ship the release left out:

```hcl
pipeline "deploy-with-test-data" {
  desc = "Loads masked test data, then deploys"
  when = "promote"

  step "INIT" {
    init "Init job home" {}
  }

  step "PRE" {
    load_items "Load job items" {}
    clone_repos "Clone job repositories" {}
  }

  step "RUN" {
    each_project "For each project in the job" {
      ship "Copy the data scripts" {
        host             = generic_server.data-runner
        local_mode       = "local_files"
        from             = "${job_dir}/${project}/ci/*.sh"
        exist_mode_local = "fail"
        to               = "/srv/deploy/${job_name}/${project}/"
        forward_only     = true
      }

      sh "Clear the last dataset" {
        host         = generic_server.data-runner
        run          = <<-EOT
          cd /srv/deploy/${job_name}/${project} && \
            sh clear-dataset.sh ${bl}
        EOT
        forward_only = true
      }

      sh "Provision test data" {
        host         = generic_server.data-runner
        run          = <<-EOT
          curl -fsS "$(cat ${gigantics_url_file})"
        EOT
        forward_only = true
      }

      sh "Check the dataset landed" {
        host         = generic_server.data-runner
        run          = <<-EOT
          cd /srv/deploy/${job_name}/${project} && \
            sh check-dataset.sh ${bl}
        EOT
        forward_only = true
        timeout      = 1200
      }
    }
  }
}
```

`${bl}` is the [environment](https://docs.clarive.com/en/concepts/environment/) the job is deploying to, under the name Clarive uses for environments internally. `gigantics_url_file` takes a different value in each environment, so the same step reads the QA pipeline's address when a [card is dropped](https://docs.clarive.com/en/getting-started/kanban/) on QA (the quality assurance environment) and the UAT pipeline's when it is dropped on UAT, where user acceptance testing happens. `ship` copies the two scripts from the application's repository to the runner. `forward_only` keeps all four steps out of a rollback, which has no use for fresh test data. The [Clarive job log](https://docs.clarive.com/en/getting-started/job-log/) shows each step under its name, so the log reads as a record of what happened to the data.

The check step's `timeout` is in seconds. Clarive cuts a command on an SSH connection off after 60 seconds unless the step sets a timeout of its own, and the check has to outlast its own polling.

![Diagram: a Clarive pipeline rule step, sh Provision test data, whose run command is curl -fsS on the contents of the file named by gigantics_url_file, branches to DEV, QA and UAT rows. Each row names that environment's URL file on the runner, its Gigantics pipeline and its sink database; the QA row is highlighted as the current job.](https://clarive.com/images/blog/where-the-test-data-comes-from/one-step-per-environment.webp)

The same rule step reaches a different pipeline and database depending on which environment the job targets.

The clear and check steps are there so the tests do not have to trust the URL. Calling it starts a job in Gigantics, and the call can return before that job has loaded anything. A check that only looked for rows would then pass right away, against whatever the previous run left behind. Emptying the tables the suite depends on first removes that false pass. The check then polls those tables until their row counts reach a minimum and stop changing, and it fails the deploy if that hasn't happened by the time it times out.

Without it, a pipeline can go green after running its whole test suite against an empty schema. That is very easy to do by accident.

The URL stays on the runner because the pipeline's API key travels inside it. By default the Clarive job log keeps the [full command line](https://docs.clarive.com/en/palette/job/run-remote/) of every step it runs on a server, with variables already filled in, so a URL kept in a Clarive variable would be written into the log of every deploy. The step passes only the file's path, and the shell on the runner reads the file. We accept one weakness in that. Anyone who can log in to the runner as the account Clarive uses can read every URL on it.

## What changes

We haven't put a percentage on any of this. Our numbers would not match another company's, and the changes are mechanical enough for any team to check against its own systems.

For testers, the environment stops being something to guard. When refilling it takes a whole Sunday, teams protect it. They nurse it back by hand after a bad run and hold meetings about who gets it next week. Refill it on every deploy and [each team can have its own](https://clarive.com/blog/why-environment-provisioning), and one that a bad test run has corrupted can be rebuilt by running the job again. The rule's Drop load option deletes the existing tables and their data before loading, so nothing from the previous run survives.

Security and compliance teams lose the standing copy, along with a question nobody could answer. Discovery reports are signed. The rule records which fields were masked and how. And the Clarive job log shows the provisioning step running for a specific release in a specific environment, with a timestamp.

For whoever owns the release, a deploy that cannot get its data now fails at the start, in the job log. Before, it could pass a test suite that never really ran.

## What is still missing

We are early with this, and several parts are more manual than we would like.

The biggest gap is knowing when the data is ready. Calling the pipeline starts a Gigantics job, and the check step infers that the job has finished from row counts that have stopped changing. That is a heuristic, and we don't know yet how well it holds up on every kind of schema. A step that waited on the Gigantics job itself, and failed with that job's own error, would be better.

Choosing the subset still takes judgment. Where and Limit can express any slice of production, but somebody has to decide which slice looks like the real thing. That person needs to know that the customer with 400 orders matters and write the rule so that customer is kept.

Reviewing datasets is manual, too. Gigantics can export the rule configuration a job used as YAML, and we want that file stored in the application's repository, next to the HCL file that holds the Clarive rule. Changing what the test data means would then go through a pull request, which records the diff and who made it and when. For now, keeping the two in step means copying the file over by hand.

The goal we have set ourselves is to delete the copy of production in staging. Securing it or masking it where it sits does not count. If the copy is still there a year from now, the project did not happen, whatever else got built along the way.

Gigantics's [documentation](https://docs.gigantics.io/en/home) covers rules and pipelines in more detail. Teams that want to work out where a step like this would fit in their own pipelines can [get in touch with us](https://clarive.com/contact).

[See what Gigantics does](https://www.gigantics.io/) 

_Disclosure: I'm also a co-founder of Gigantics._
