For the complete documentation index, see llms.txt. This page is also available as Markdown.

Before you create a MySQL workspace

Before you create the workspace, the source and destination databases should already exist.

Database setup requirements

Source and destination on different servers

When Tonic Structural generates data for MySQL, the destination database uses the same name as the source database.

To prevent a conflict, the source and destination databases must be on different servers.

Matching MySQL tablespaces on source and destination

The MySQL tablespaces on the source database must exist in the destination database.

Enable local file loading on the destination database

Structural writes data from the source database to files on the destination database. It then uses the LOAD DATA statement to load the files into the destination database. If LOAD DATA is not enabled on the destination database, then data generation fails.

To enable LOAD DATA, run the following command on the destination database:

SET GLOBAL local_infile = 'ON';

Creating database users for Structural

On each database, you must create a user that has the required permissions that Structural needs to function.

Creating the source database user

The following is an example of how to create a new user, called tonic, and then grant the necessary permissions.

For the source database, we recommend that you use a backup or fast follower database instead of a direct connection to your production environment.

If you have stored routines that other database objects reference, then you must grant permissions for routines. Otherwise your jobs will fail. Stored routines include procedures and functions.

If you have triggers or events that are important to the functionality of your database, then you should also grant permissions for triggers or events.

When you specify a GRANT options for an object type, then Structural copies that object type from the source to the destination database. Otherwise the object type is excluded.

To verify the granted permissions, run show grants for tonic. The output is something like:

Creating the destination database user

Configuring how Structural reads MySQL source data

The first step in data generation is to read the source data. Structural only needs to read source data from tables that use the De-Identify or Incremental table modes.

For Incremental mode, Structural always reads each table straight through.

For De-identify mode, how Structural reads the data is based on the size of the table and the Structural configuration.

Using a single query per table

By default, batched reads are disabled, and Structural uses a single query to read each table.

Enabling and configuring batched reads

When batched reads are enabled, Structural reads the table in batches that are sized by data volume. The rows stream in primary-key order. When a batch reaches a target amount of data, the batch ends. Batched reads replace a single long query with a series of shorter ones.

To enable batched reads, set the environment setting TONIC_MYSQL_KEYSET_BATCH_TARGET_MEGABYTES to the number of megabytes in each batch. You can configure this setting from the Environment Settings tab on Structural Settings. The default value is 0, which indicates that batched reads are disabled.

When batched reads are enabled, the environment setting TONIC_MYSQL_KEYSET_MAX_ROWS_PER_BATCH also sets a limit on the number of rows. You can also configure this setting from the Environment Settings tab on Structural Settings.

For example, when data is sparse or a table is narrow, it might take a long time for a batch to reach the data volume limit. When the batch reaches the maximum number of rows, then even if it has not yet reached the data volume limit, Structural starts a new batch. By default, the row limit is 500,000.

Batched reads require a primary key. If the data does not have a primary key, then even if batched reads are enabled, Structural uses the single query.

Configuring whether Structural creates the destination database schema

If you provide custom names for destination database schemas, then you cannot create the schemas yourself.

By default, during each data generation job, Structural creates the database schema for the destination database tables, then populates the database tables based on the workspace configuration.

If you prefer to manage the destination database schema yourself, then set the environment setting TONIC_MYSQL_SKIP_CREATE_DB to true. You can configure TONIC_MYSQL_SKIP_CREATE_DB from the Environment Settings tab on Structural Settings.

When TONIC_MYSQL_SKIP_CREATE_DB is true, then Structural does not create the destination database schema. Before you run data generation, you must create the destination database with the full schema.

Note that you can override this setting in individual workspaces. For more information, go to Advanced workspace overrides.

During data generation, Structural deletes the data from the destination database tables, except in the following cases:

  • Tables that use Preserve Destination or Incremental mode

  • Upsert data generation

It then populates the tables with the new destination data.

For a diagram of the data generation process when you manage the destination schema, go to Data process - User-managed destination schema.

Last updated

Was this helpful?