Skip to content

About

Terraform module: terraform-google-datastream-stream

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

☁️ Google Cloud Datastream Stream Terraform Module

Provisions a single Datastream CDC replication stream (google_datastream_stream) tying a source connection profile, a destination connection profile, object selection, and a backfill strategy into one coherent unit. Targets hashicorp/google ~> 7.0, Terraform >= 1.12.0.

Terraform Provider Version Type Resources Posture


🧩 Overview

  • 🔀 Provisions one Datastream stream (google_datastream_stream.this) — the terminal resource of a Datastream CDC composition. It never creates a connection profile itself; it always references TWO separate, already-existing terraform-google-datastream-connection-profile instances by id (one source, one destination).
  • 🗄️ v1 scopes per-engine source modeling to MySQL, Oracle, PostgreSQL, and SQL Server, each with full database/schema → table → column include/exclude granularity. salesforce_source_config/spanner_source_config/mongodb_source_config (and their backfill_all exclusion-tree counterparts) and the entire rule_sets argument are OUT OF SCOPE for v1 — deliberate scope boundaries, not oversights (see 🧠 Architecture Notes).
  • 🎯 Destination modeling covers both gcs_destination_config and bigquery_destination_config in full, including BigQuery's dataset-targeting oneof (single_target_dataset vs. source_hierarchy_datasets) and write-mode oneof (merge vs. append_only).
  • ⚠️ The single most important operational fact about this module: nothing in this library can verify at plan time that the source-engine block set here (e.g. mysql_source_config) actually matches the engine of the connection profile referenced by source_connection_profile. A mismatch passes validate/plan cleanly and fails only at apply, surfacing as state = "FAILED". See 🔍 Troubleshooting.
  • 🚦 backfill_all vs. backfill_none is modeled as a mandatory, explicit choice — a validation {} block requires exactly one, rather than letting an ambiguous "neither set" configuration plan successfully.
  • 🛑 desired_state is left unset by default (resolving to the provider's own "NOT_STARTED") — the safe, inert result. A stream this module creates does not replicate anything until the caller explicitly sets desired_state = "RUNNING".

💡 Why it matters: a Datastream stream is where a CDC pipeline either works or silently fails — get the source-engine/connection-profile-type pairing and the backfill strategy right at authoring time, since this library's plan-only posture cannot catch either mistake before apply.


❤️ Support this project

If these Terraform modules have been helpful to you or your organization, I'd appreciate your support in any of the following ways:

Whether it's a star, a professional connection, or a coffee, every gesture helps keep these modules actively maintained and continually improving. Thank you for being part of the community!


🗺️ Where this fits

This module has no downstream consumer identified yet in the initial catalog — it is a leaf/terminal module, the end of a Datastream CDC composition. Its upstream inputs are two mandatory connection-profile ids (source and destination, from two separate terraform-google-datastream-connection-profile instances) plus two independently optional inputs: a BigQuery dataset id (only relevant for a BigQuery destination using single_target_dataset) and a KMS crypto key id, consumed at TWO distinct points (the stream's own customer_managed_encryption_key, and separately, bigquery_destination_config.source_hierarchy_datasets.dataset_template.kms_key_name).

flowchart LR
 src["terraform-google-datastream-connection-profile<br/>(source instance)"]:::sibling
 dst["terraform-google-datastream-connection-profile<br/>(destination instance)"]:::sibling
 bq["terraform-google-bigquery-dataset"]:::sibling
 kms["terraform-google-kms-keyring"]:::sibling
 stream["terraform-google-datastream-stream"]:::self

 src -->|"id (source profile)"| stream
 dst -->|"id (destination profile)"| stream
 bq -->|"dataset_id (optional, BigQuery destination)"| stream
 kms -->|"crypto_key_id (optional, stream-level CMEK)"| stream
 kms -->|"crypto_key_id (optional, dataset_template CMEK)"| stream

 classDef self fill:#4285F4,color:#FFFFFF,stroke:#174EA6,stroke-width:2px;
 classDef sibling fill:#ECEFF1,color:#263238,stroke:#90A4AE,stroke-width:1px;
Loading

Validated via the Mermaid Chart MCP (validate_and_render_mermaid_diagram) before embedding.


🧬 What this builds

flowchart TB
 subgraph Module["terraform-google-datastream-stream"]
 direction TB
 this["google_datastream_stream.this<br/>(keystone)"]:::keystone
 sc["source_config<br/>(required, exactly one engine)"]:::nested
 dc["destination_config<br/>(required, exactly one type)"]:::nested
 ba["backfill_all / backfill_none<br/>(exactly one)"]:::nested

 subgraph InScope["v1 in-scope source engines"]
 direction LR
 mysql["mysql_source_config"]:::inscope
 oracle["oracle_source_config"]:::inscope
 pg["postgresql_source_config"]:::inscope
 sqls["sql_server_source_config"]:::inscope
 end

 subgraph OutScope["v1 out-of-scope source engines"]
 direction LR
 sf["salesforce_source_config"]:::outscope
 spn["spanner_source_config"]:::outscope
 mongo["mongodb_source_config"]:::outscope
 end

 gcs["gcs_destination_config"]:::inscope
 bq["bigquery_destination_config"]:::inscope
 end

 this --> sc
 this --> dc
 this --> ba
 sc --> InScope
 sc -.->|"out of scope for v1"| OutScope
 dc --> gcs
 dc --> bq

 classDef keystone fill:#174EA6,color:#FFFFFF,stroke:#0D47A1,stroke-width:2px;
 classDef nested fill:#ECEFF1,color:#263238,stroke:#90A4AE,stroke-width:1px;
 classDef inscope fill:#4285F4,color:#FFFFFF,stroke:#174EA6,stroke-width:2px;
 classDef outscope fill:#CFD8DC,color:#546E7A,stroke:#90A4AE,stroke-width:1px,stroke-dasharray: 3 3;
Loading

Validated via the Mermaid Chart MCP before embedding.

Resource inventory (1 resource):

Resource Count Role
google_datastream_stream.this 1 Keystone — the entire module is this one resource; source_config/destination_config/backfill_all are internal structure, not separate managed resources

✅ Provider / Versions

Requirement Value
Terraform >= 1.12.0
hashicorp/google ~> 7.0 (resolved 7.39.0 at authoring time)
Provider block None — the caller's root module configures google (ADC, WIF, or a service-account key per our authentication model)

Schema notes that bite:

  • No self_link attribute exists anywhere on this resource. Confirmed absent from the live schema — only id, name, state, terraform_labels, effective_labels are computed. state fills the house id-then-self_link output slot instead (see 🧾 Outputs).
  • The engine-mismatch apply-time-only failure mode. source_config's per-engine block (mysql_source_config etc.) is a SIBLING optional field to source_connection_profile, not type-derived from it. A caller who points source_connection_profile at a PostgreSQL-type profile but supplies mysql_source_config produces a configuration that PASSES validate/plan cleanly and FAILS ONLY at apply — see 🔍 Troubleshooting (the first, most prominent entry) and 🧪 Testing.
  • backfill_all vs. backfill_none is a real, consequential, non-cosmetic replication-scope choice, enforced here via a mandatory validation {} requiring exactly one.
  • desired_state defaults to "NOT_STARTED". A stream created with no explicit desired_state is fully configured but replicates nothing until the caller flips it to "RUNNING" — see Example 8.
  • gcs_destination_config has NO bucket argument. The destination bucket comes entirely from whichever gcs_profile-type connection profile is referenced via destination_connection_profile; path is only a sub-path/prefix inside that bucket.
  • postgresql_source_config has no max_concurrent_cdc_tasks field — unlike mysql/oracle/sql_server, which all have it — a confirmed, real per-engine schema asymmetry.
  • Correction found during this module's own authoring session: postgresql_source_config. replication_slot and .publication are BOTH REQUIRED once that block is set (confirmed against the live schema; this module's original scaffolding brief described them as optional pass-throughs — the live schema is ground truth).
  • sql_server_source_config and oracle_source_config columns expose ONLY column/data_type as caller-settable — length/nullable/ordinal_position/precision/primary_key/scale/ encoding are all Computed-only on the live schema for these two engines. mysql_source_config and postgresql_source_config columns expose a wider settable set (see 📥 Inputs). Do not assume parity across engines.
  • A live provider-doc quirk: the rendered Argument Reference for oracle_source_config is itself mislabeled "MySQL data source configuration." in the live 7.39.0 docs — a copy-paste artifact in the provider's own description string, confirmed in both the schema JSON and the live provider documentation's chunked output. Not a mistake in this module.
  • v1 scope exclusions: salesforce_source_config/spanner_source_config/ mongodb_source_config (and their backfill_all counterparts) and the entire rule_sets argument exist on the live resource but are out of scope for this module's v1 — see 🧠 Architecture Notes for the full reasoning.
  • min_items: 1 on every engine's include_objects/exclude_objects database/schema list. Confirmed in the live schema (mysql_databases, oracle_schemas, postgresql_schemas, sql server's schemas all require at least one entry once the parent block is present). This module does not duplicate that constraint in its own validation {} — Terraform core enforces a resource's own nested-block min_items natively once the dynamic blocks render, surfacing as an "Insufficient <block> blocks" error at terraform validate time.

🔑 Required IAM Roles

  • roles/datastream.admin — least-privilege role for the applying principal; matches the sibling terraform-google-datastream-connection-profile module's role.

(Sourced directly from SCOPE.md — not re-derived.)


☁️ GCP Prerequisites

  • datastream.googleapis.com must already be enabled on the target project (via terraform-google-project-services) before this module applies.
  • Both connection profiles this stream references (source_connection_profile, destination_connection_profile) must already exist — this module consumes their ids by reference, it does not create them.
  • For a PostgreSQL source, the logical replication slot (replication_slot) and publication (publication) named in postgresql_source_config must already exist on the source database — external prerequisites this module consumes by reference, confirmed REQUIRED once postgresql_source_config is set.
  • If CMEK is used, the Datastream/BigQuery service agent for the target project requires an IAM grant on the referenced crypto key before encryption succeeds at apply — out of scope for this module by design, subject to GCP's ~60-second IAM propagation delay.

(Sourced directly from SCOPE.md — not re-derived.)


📁 Module Structure

terraform-google-datastream-stream/
├── providers.tf # required_providers (hashicorp/google ~> 7.0) + required_version — no provider {} block
├── variables.tf # display_name, stream_id, location, source_config, destination_config, backfill_all/backfill_none, labels, timeouts
├── main.tf # google_datastream_stream.this
├── outputs.tf # id, state, name
├── README.md # this file
├── SCOPE.md # cross-module contract
└── examples/ # runnable example(s) matching the Quick Start below

⚙️ Quick Start

# Caller's root module configures the google provider (ADC, WIF, or a service-account key) —
# this module never declares project/region/zone/credentials variables. Also assumes two
# already-applied terraform-google-datastream-connection-profile instances (source + destination) —
# see 🏗️ End-to-end composition for the full picture.

module "orders_cdc_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "orders-mysql-to-gcs"
  stream_id    = "orders-mysql-to-gcs"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.mysql_source_profile.id
    mysql_source_config = {
      include_objects = {
        mysql_databases = [
          { database = "orders_db" }
        ]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.gcs_destination_profile.id
    gcs_destination_config = {
      path             = "cdc/orders/"
      avro_file_format = true
    }
  }

  # A deliberate choice, not a default — see backfill_all/backfill_none in 🧱 Design Principles.
  backfill_none = true
}

🔌 Cross-Module Contract

Consumes

Input Type Source module
Source connection profile id string (required) terraform-google-datastream-connection-profile (source-type instance)
Destination connection profile id string (required) terraform-google-datastream-connection-profile (destination-type instance, SEPARATE from the source instance)
BigQuery dataset id (projects/{project}/datasets/{dataset_id} or {project}:{dataset_id}) string (optional) terraform-google-bigquery-dataset
KMS crypto key id/name (two independent consumption points: stream-level customer_managed_encryption_key, destination-dataset-level dataset_template.kms_key_name) string (optional) terraform-google-kms-keyring

Emits

Output Description Consumed by
id Stream resource id (projects/{{project}}/locations/{{location}}/streams/{{stream_id}}) none identified yet — leaf/terminal module
state API-reported actual state — replaces self_link (confirmed absent); distinct from var.desired_state none identified yet
name Fully-qualified stream name as reported by the API none identified yet

📚 Example Library

1 · Minimal MySQL-to-GCS stream, backfill_none
module "orders_cdc_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "orders-mysql-to-gcs"
  stream_id    = "orders-mysql-to-gcs"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.mysql_source_profile.id
    mysql_source_config = {
      include_objects = {
        mysql_databases = [
          { database = "orders_db" }
        ]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.gcs_destination_profile.id
    gcs_destination_config = {
      path             = "cdc/orders/"
      avro_file_format = true
    }
  }

  backfill_none = true
}

ℹ️ The shallowest correct composition this module supports — backfill_none means only new changes from this point forward are replicated, no historical backfill.

2 · MySQL-to-GCS with backfill_all and specific mysql_excluded_objects
module "customers_cdc_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "customers-mysql-to-gcs"
  stream_id    = "customers-mysql-to-gcs"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.mysql_source_profile.id
    mysql_source_config = {
      include_objects = {
        mysql_databases = [
          { database = "customers_db" }
        ]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.gcs_destination_profile.id
    gcs_destination_config = {
      path             = "cdc/customers/"
      avro_file_format = true
    }
  }

  backfill_all = {
    mysql_excluded_objects = {
      mysql_databases = [
        {
          database = "customers_db"
          mysql_tables = [
            { table = "audit_log" }
          ]
        }
      ]
    }
  }
}

⚠️ backfill_all replicates ALL historical data for every object in include_objects (here, everything in customers_db) EXCEPT what's listed in mysql_excluded_objects — here, the audit_log table. Consider data volume before choosing this over backfill_none (see 🔍 Troubleshooting).

3 · PostgreSQL-to-BigQuery using single_target_dataset
module "pg_orders_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "pg-orders-to-bq"
  stream_id    = "pg-orders-to-bq"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.postgresql_source_profile.id
    postgresql_source_config = {
      replication_slot = "datastream_orders_slot"
      publication      = "datastream_orders_pub"
      include_objects = {
        postgresql_schemas = [
          { schema = "public" }
        ]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.bq_destination_profile.id
    bigquery_destination_config = {
      single_target_dataset = {
        dataset_id = module.orders_dataset.id
      }
      merge = true
    }
  }

  backfill_none = true
}

💡 single_target_dataset.dataset_id consumes terraform-google-bigquery-dataset's id output directly. merge is BigQuery's default write mode (all changes merged into current state) — append_only is the alternative (see Example 14).

4 · PostgreSQL source with replication_slot/publication and object exclusions
module "pg_inventory_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "pg-inventory-to-gcs"
  stream_id    = "pg-inventory-to-gcs"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.postgresql_source_profile.id
    postgresql_source_config = {
      replication_slot = "datastream_inventory_slot"
      publication      = "datastream_inventory_pub"
      include_objects = {
        postgresql_schemas = [
          { schema = "inventory" }
        ]
      }
      exclude_objects = {
        postgresql_schemas = [
          {
            schema = "inventory"
            postgresql_tables = [
              { table = "staging_temp" }
            ]
          }
        ]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.gcs_destination_profile.id
    gcs_destination_config = {
      path = "cdc/inventory/"
      json_file_format = {
        schema_file_format = "NO_SCHEMA_FILE"
        compression        = "GZIP"
      }
    }
  }

  backfill_none = true
}

⚠️ replication_slot and publication are BOTH REQUIRED once postgresql_source_config is set, and — confirmed against the live schema — they are NOT created by this module or by terraform-google-datastream-connection-profile. The named logical replication slot and publication must already exist on the source PostgreSQL database before this stream can start replicating; see ☁️ GCP Prerequisites and 🔍 Troubleshooting.

5 · Oracle-to-BigQuery stream
module "oracle_finance_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "oracle-finance-to-bq"
  stream_id    = "oracle-finance-to-bq"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.oracle_source_profile.id
    oracle_source_config = {
      include_objects = {
        oracle_schemas = [
          { schema = "FINANCE" }
        ]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.bq_destination_profile.id
    bigquery_destination_config = {
      single_target_dataset = {
        dataset_id = module.finance_dataset.id
      }
      merge = true
    }
  }

  backfill_none = true
}

ℹ️ oracle_source_config columns expose only column/data_type as caller-settable — precision/scale/encoding/etc. are Computed-only, populated by the API on read.

6 · SQL Server-to-GCS stream with change_tables CDC method
module "sqlserver_hr_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "sqlserver-hr-to-gcs"
  stream_id    = "sqlserver-hr-to-gcs"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.sqlserver_source_profile.id
    sql_server_source_config = {
      change_tables = true
      include_objects = {
        schemas = [
          { schema = "hr" }
        ]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.gcs_destination_profile.id
    gcs_destination_config = {
      path             = "cdc/hr/"
      avro_file_format = true
    }
  }

  backfill_none = true
}

💡 change_tables and transaction_logs are the two alternative CDC read methods for SQL Server — mutually exclusive, enforced by this module's validation {}. Omit both to let the API's own default apply.

7 · BigQuery destination using source_hierarchy_datasets with its own kms_key_name
module "multi_schema_bq_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "oracle-multi-schema-to-bq"
  stream_id    = "oracle-multi-schema-to-bq"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.oracle_source_profile.id
    oracle_source_config = {
      include_objects = {
        oracle_schemas = [
          { schema = "SALES" },
          { schema = "MARKETING" },
        ]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.bq_destination_profile.id
    bigquery_destination_config = {
      source_hierarchy_datasets = {
        dataset_template = {
          location          = "us-east1"
          dataset_id_prefix = "oracle_cdc"
          # A SEPARATE, independently-optional CMEK reference from the stream-level
          # customer_managed_encryption_key — scoped only to datasets created by this template.
          kms_key_name = module.kms_keyring.crypto_key_ids["bq-dataset-key"]
        }
      }
      merge = true
    }
  }

  backfill_none = true
}

🔒 dataset_template.kms_key_name is a SEPARATE CMEK reference from the stream's own customer_managed_encryption_key (Example 9) — do not conflate the two. Both are optional and never defaulted to a specific key by this module.

8 · Stream created with desired_state = NOT_STARTED, then flipped to RUNNING
module "staged_mysql_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "staged-mysql-to-gcs"
  stream_id    = "staged-mysql-to-gcs"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.mysql_source_profile.id
    mysql_source_config = {
      include_objects = {
        mysql_databases = [
          { database = "staging_db" }
        ]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.gcs_destination_profile.id
    gcs_destination_config = {
      path             = "cdc/staging/"
      avro_file_format = true
    }
  }

  backfill_none = true

  # Explicit and deliberate — fully configured, but replicates nothing yet.
  desired_state = "NOT_STARTED"
}

Once ready to begin replication, a follow-up apply flips only desired_state:

desired_state = "RUNNING"

ℹ️ This is a genuine two-step operational pattern, not a workaround: create the stream inert (NOT_STARTED, this module's own default when desired_state is omitted entirely), validate the configuration and source connectivity out of band, then flip to RUNNING in a separate, reviewed apply. Check this module's state output (not desired_state) to confirm the API actually reports RUNNING afterward.

9 · Stream-level customer_managed_encryption_key set
module "encrypted_mysql_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "encrypted-mysql-to-gcs"
  stream_id    = "encrypted-mysql-to-gcs"
  location     = "us-east1"

  customer_managed_encryption_key = module.kms_keyring.crypto_key_ids["datastream-key"]

  source_config = {
    source_connection_profile = module.mysql_source_profile.id
    mysql_source_config = {
      include_objects = {
        mysql_databases = [
          { database = "sensitive_db" }
        ]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.gcs_destination_profile.id
    gcs_destination_config = {
      path             = "cdc/sensitive/"
      avro_file_format = true
    }
  }

  backfill_none = true
}

🔒 Never defaulted to a specific key by this module — consumed from terraform-google-kms-keyring's crypto key id output. The Datastream service agent needs an IAM grant on this key before it succeeds at apply (out of scope here; see ☁️ GCP Prerequisites).

10 · GCS destination — json_file_format vs. avro_file_format
# Variant A: JSON output, no embedded schema, gzip-compressed.
module "json_output_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "orders-mysql-to-gcs-json"
  stream_id    = "orders-mysql-to-gcs-json"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.mysql_source_profile.id
    mysql_source_config = {
      include_objects = {
        mysql_databases = [{ database = "orders_db" }]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.gcs_destination_profile.id
    gcs_destination_config = {
      path = "cdc/orders-json/"
      json_file_format = {
        schema_file_format = "NO_SCHEMA_FILE"
        compression        = "GZIP"
      }
    }
  }

  backfill_none = true
}

# Variant B: AVRO output (self-describing schema, this module's Quick Start default).
module "avro_output_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "orders-mysql-to-gcs-avro"
  stream_id    = "orders-mysql-to-gcs-avro"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.mysql_source_profile.id
    mysql_source_config = {
      include_objects = {
        mysql_databases = [{ database = "orders_db" }]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.gcs_destination_profile.id
    gcs_destination_config = {
      path             = "cdc/orders-avro/"
      avro_file_format = true
    }
  }

  backfill_none = true
}

⚠️ avro_file_format and json_file_format are mutually exclusive, enforced by this module's validation {}. Neither is Required by the live schema — leaving both unset may fail at apply (invisible to this library's plan-only gate), so pick one deliberately.

11 · MySQL source using gtid CDC method (instead of binary_log_position)
module "gtid_mysql_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "gtid-mysql-to-gcs"
  stream_id    = "gtid-mysql-to-gcs"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.mysql_source_profile.id
    mysql_source_config = {
      gtid = true
      include_objects = {
        mysql_databases = [
          { database = "orders_db" }
        ]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.gcs_destination_profile.id
    gcs_destination_config = {
      path             = "cdc/orders-gtid/"
      avro_file_format = true
    }
  }

  backfill_none = true
}

ℹ️ gtid and binary_log_position are mutually exclusive CDC read methods for MySQL — this requires the source MySQL server actually have GTID-based replication enabled; omit both to let the API apply its own default.

12 · Oracle source with drop_large_objects LOB handling
module "oracle_docs_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "oracle-docs-to-gcs"
  stream_id    = "oracle-docs-to-gcs"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.oracle_source_profile.id
    oracle_source_config = {
      drop_large_objects = true
      include_objects = {
        oracle_schemas = [
          { schema = "DOCS" }
        ]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.gcs_destination_profile.id
    gcs_destination_config = {
      path             = "cdc/docs/"
      avro_file_format = true
    }
  }

  backfill_none = true
}

⚠️ drop_large_objects and stream_large_objects are alternative, mutually exclusive LOB-handling strategies. Dropping large objects (BLOB/CLOB-shaped columns) reduces replication overhead but means that data is NOT replicated at all — confirm this is acceptable before choosing it over stream_large_objects.

13 · SQL Server source with column-level include_objects granularity
module "sqlserver_payroll_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "sqlserver-payroll-to-gcs"
  stream_id    = "sqlserver-payroll-to-gcs"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.sqlserver_source_profile.id
    sql_server_source_config = {
      transaction_logs = true
      include_objects = {
        schemas = [
          {
            schema = "payroll"
            tables = [
              {
                table = "employee_pay"
                columns = [
                  { column = "employee_id" },
                  { column = "pay_period" },
                  { column = "gross_pay" },
                ]
              }
            ]
          }
        ]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.gcs_destination_profile.id
    gcs_destination_config = {
      path             = "cdc/payroll/"
      avro_file_format = true
    }
  }

  backfill_none = true
}

ℹ️ Naming specific columns under a table restricts replication to only those columns. column/data_type are the only caller-settable fields for SQL Server columns (same narrow set as Oracle) — length/nullable/ordinal_position/precision/primary_key/scale are all Computed-only and cannot be set here.

14 · BigQuery destination — merge vs. append_only write mode
# Variant A: merge (BigQuery's own default write mode — reflects current source state, no history).
module "merge_mode_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "orders-pg-to-bq-merge"
  stream_id    = "orders-pg-to-bq-merge"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.postgresql_source_profile.id
    postgresql_source_config = {
      replication_slot = "datastream_orders_slot"
      publication      = "datastream_orders_pub"
      include_objects = {
        postgresql_schemas = [{ schema = "public" }]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.bq_destination_profile.id
    bigquery_destination_config = {
      single_target_dataset = { dataset_id = module.orders_dataset.id }
      merge                 = true
    }
  }

  backfill_none = true
}

# Variant B: append_only (retains full historical change events — INSERT/UPDATE-INSERT/
# UPDATE-DELETE/DELETE — as distinct rows).
module "append_only_mode_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "orders-pg-to-bq-append"
  stream_id    = "orders-pg-to-bq-append"
  location     = "us-east1"

  source_config = {
    source_connection_profile = module.postgresql_source_profile.id
    postgresql_source_config = {
      replication_slot = "datastream_orders_audit_slot"
      publication      = "datastream_orders_audit_pub"
      include_objects = {
        postgresql_schemas = [{ schema = "public" }]
      }
    }
  }

  destination_config = {
    destination_connection_profile = module.bq_destination_profile.id
    bigquery_destination_config = {
      single_target_dataset = { dataset_id = module.orders_audit_dataset.id }
      append_only           = true
    }
  }

  backfill_none = true
}

💡 merge and append_only are mutually exclusive and this module requires exactly one be set — a deliberate choice, since which one a caller needs depends entirely on whether downstream analytics need historical change events (append_only) or only current state (merge).

15 · 🏗️ End-to-end composition

Wires two separate terraform-google-datastream-connection-profile instances (source + destination) into this module's source_config/destination_config, plus terraform-google-bigquery-dataset's id for the BigQuery destination — reflecting every relationship documented in this module's SCOPE.md.

# Source: an existing MySQL server, referenced by a source-type connection profile.
module "mysql_source_profile" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-connection-profile.git?ref=v1.0.0"

  display_name          = "orders-mysql-source"
  connection_profile_id = "orders-mysql-source"
  location              = "us-east1"

  mysql_profile = {
    hostname                       = "10.20.0.10"
    username                       = "datastream_reader"
    secret_manager_stored_password = module.mysql_reader_secret.secret_version_id
  }
}

# Destination: a BigQuery-type connection profile.
module "bq_destination_profile" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-connection-profile.git?ref=v1.0.0"

  display_name          = "orders-bq-destination"
  connection_profile_id = "orders-bq-destination"
  location              = "us-east1"

  bigquery_profile = {}
}

# The BigQuery dataset this stream's single_target_dataset targets.
module "orders_dataset" {
  source = "git::https://github.com/microsoftexpert/terraform-google-bigquery-dataset.git?ref=v1.0.0"

  dataset_id = "casey_orders_cdc"
  location   = "us-east1"
}

# The stream itself — this module.
module "orders_cdc_stream" {
  source = "git::https://github.com/microsoftexpert/terraform-google-datastream-stream.git?ref=v1.0.0"

  display_name = "orders-mysql-to-bq"
  stream_id    = "orders-mysql-to-bq"
  location     = "us-east1"

  source_config = {
    # Consumes terraform-google-datastream-connection-profile's id output (source instance).
    source_connection_profile = module.mysql_source_profile.id
    mysql_source_config = {
      include_objects = {
        mysql_databases = [
          { database = "orders_db" }
        ]
      }
    }
  }

  destination_config = {
    # Consumes terraform-google-datastream-connection-profile's id output (destination instance —
    # a SEPARATE instance from the source above).
    destination_connection_profile = module.bq_destination_profile.id
    bigquery_destination_config = {
      # Consumes terraform-google-bigquery-dataset's id output.
      single_target_dataset = {
        dataset_id = module.orders_dataset.id
      }
      merge = true
    }
  }

  backfill_all = {
    mysql_excluded_objects = {
      mysql_databases = [
        {
          database = "orders_db"
          mysql_tables = [
            { table = "session_cache" }
          ]
        }
      ]
    }
  }

  desired_state = "RUNNING"
}

⚠️ Before this composition can succeed, the ENGINE of mysql_source_config must actually match the type of module.mysql_source_profile (a mysql_profile-type connection profile) — this library's plan-only posture cannot verify that pairing; a mismatch passes validate/plan and fails only at apply. Also confirm datastream.googleapis.com and bigquery.googleapis.com are already enabled (via terraform-google-project-services) and both connection profiles have already been applied before this stream's own apply.


📥 Inputs

Variable Type Default Notes
display_name string — (required)
stream_id string — (required) Force-new
location string — (required) Force-new; not validated against a hardcoded list
source_config object({...}) — (required) Exactly one per-engine block (mysql/oracle/postgresql/sql_server); see full schema below
destination_config object({...}) — (required) Exactly one of gcs/bigquery; see full schema below
backfill_all object({...}) null Exactly one of backfill_all/backfill_none required
backfill_none bool false Exactly one of backfill_all/backfill_none required
customer_managed_encryption_key string null Never defaulted to a specific key
create_without_validation bool false Skips Datastream's own live connectivity validation
deletion_policy string "PREVENT" House extension — resource has no deletion_protection
desired_state string null (→ "NOT_STARTED") Deliberately inert by default
labels map(string) {} GCP label format validated
timeouts object({create,update,delete}) null 20-minute provider default each
Full source_config object schema
variable "source_config" {
 type = object({
 source_connection_profile = string

 mysql_source_config = optional(object({
 max_concurrent_cdc_tasks = optional(number)
 max_concurrent_backfill_tasks = optional(number)
 binary_log_position = optional(bool, false)
 gtid = optional(bool, false)
 include_objects = optional(object({
 mysql_databases = list(object({
 database = string
 mysql_tables = optional(list(object({
 table = string
 mysql_columns = optional(list(object({
 column = optional(string)
 data_type = optional(string)
 collation = optional(string)
 primary_key = optional(bool)
 nullable = optional(bool)
 ordinal_position = optional(number)
 })), [])
 })), [])
 }))
 }))
 exclude_objects = optional(object({ mysql_databases = list(object({... })) }))
 }))

 oracle_source_config = optional(object({
 max_concurrent_cdc_tasks = optional(number)
 max_concurrent_backfill_tasks = optional(number)
 drop_large_objects = optional(bool, false)
 stream_large_objects = optional(bool, false)
 include_objects = optional(object({
 oracle_schemas = list(object({
 schema = string
 oracle_tables = optional(list(object({
 table = string
 oracle_columns = optional(list(object({
 column = optional(string)
 data_type = optional(string)
 })), [])
 })), [])
 }))
 }))
 exclude_objects = optional(object({ oracle_schemas = list(object({... })) }))
 }))

 postgresql_source_config = optional(object({
 replication_slot = string # REQUIRED once this block is set
 publication = string # REQUIRED once this block is set
 max_concurrent_backfill_tasks = optional(number)
 # No max_concurrent_cdc_tasks — confirmed absent for this engine.
 include_objects = optional(object({
 postgresql_schemas = list(object({
 schema = string
 postgresql_tables = optional(list(object({
 table = string
 postgresql_columns = optional(list(object({
 column = optional(string)
 data_type = optional(string)
 primary_key = optional(bool)
 nullable = optional(bool)
 ordinal_position = optional(number)
 })), [])
 })), [])
 }))
 }))
 exclude_objects = optional(object({ postgresql_schemas = list(object({... })) }))
 }))

 sql_server_source_config = optional(object({
 max_concurrent_cdc_tasks = optional(number)
 max_concurrent_backfill_tasks = optional(number)
 transaction_logs = optional(bool, false)
 change_tables = optional(bool, false)
 include_objects = optional(object({
 schemas = list(object({
 schema = string
 tables = optional(list(object({
 table = string
 columns = optional(list(object({
 column = optional(string)
 data_type = optional(string)
 })), [])
 })), [])
 }))
 }))
 exclude_objects = optional(object({ schemas = list(object({... })) }))
 }))
 })
}
Full destination_config object schema
variable "destination_config" {
  type = object({
    destination_connection_profile = string

    gcs_destination_config = optional(object({
      path                   = optional(string)
      file_rotation_mb       = optional(number)
      file_rotation_interval = optional(string)
      avro_file_format       = optional(bool, false)
      json_file_format = optional(object({
        schema_file_format = optional(string) # NO_SCHEMA_FILE | AVRO_SCHEMA_FILE
        compression        = optional(string) # NO_COMPRESSION | GZIP
      }))
    }))

    bigquery_destination_config = optional(object({
      data_freshness = optional(string)
      single_target_dataset = optional(object({
        dataset_id = string
      }))
      source_hierarchy_datasets = optional(object({
        project_id = optional(string)
        dataset_template = object({
          dataset_id_prefix = optional(string)
          kms_key_name      = optional(string)
          location          = string
        })
      }))
      merge       = optional(bool, false)
      append_only = optional(bool, false)
      blmt_config = optional(object({
        bucket          = string
        connection_name = string
        file_format     = string
        table_format    = string
        root_path       = optional(string)
      }))
    }))
  })
}

🧾 Outputs

Output Description Sensitive
id Stream Terraform-internal resource id No
state API-reported actual state — replaces self_link No
name Fully-qualified stream name No

🧠 Architecture Notes

  • The engine-mismatch apply-time-only failure mode is the single most important operational note in this document. source_config.<engine>_source_config is a sibling optional field to source_connection_profile, not derived from its type. Terraform cannot cross-check that the engine block set here matches the actual profile type referenced — a mismatch passes validate/plan cleanly and surfaces only as state = "FAILED" on a later apply/refresh. See 🔍 Troubleshooting.
  • backfill_all vs. backfill_none is enforced as a mandatory, explicit choice via validation {} — never left to plan successfully with both unset. Choosing backfill_all on a large source can consume significant time and Datastream quota; this is a deliberate, data-volume-dependent decision the calling composition must make, not a default to accept blindly.
  • desired_state defaults to NOT_STARTED. A stream this module creates is fully configured but replicates nothing until the caller explicitly sets desired_state = "RUNNING" — see Example 8's two-step create-then-start pattern. This module's state output reflects what actually happened, which can diverge from desired_state.
  • gcs_destination_config has no bucket argument. The destination bucket is entirely determined by the referenced gcs_profile-type connection profile; path is only a sub-path/prefix inside it. A future maintainer looking for a missing bucket field here should look at the connection profile instead.
  • postgresql_source_config has no max_concurrent_cdc_tasks — a real, confirmed per-engine schema asymmetry versus mysql/oracle/sql_server.
  • replication_slot/publication are REQUIRED, external prerequisites once postgresql_source_config is set — this module does not create them, and they must already exist on the source database (a finding confirmed during this module's own authoring session, correcting an earlier assumption that they were optional).
  • Per-engine column-settable-field asymmetry. mysql_source_config/postgresql_source_config columns expose primary_key/nullable/ordinal_position as settable; oracle_source_config/ sql_server_source_config columns expose ONLY column/data_type — everything else is Computed-only on the live schema. Do not assume parity across engines when extending this module.
  • v1 scope exclusions: salesforce_source_config/spanner_source_config/ mongodb_source_config (and their backfill_all counterparts) and the entire rule_sets argument exist on the live resource but are explicitly out of scope for this module's v1 — modeling all seven engines' full column-level granularity plus rule_sets would produce one of the largest variables.tf files in this library, larger than terraform-google-gke-cluster or terraform-google-cloud-sql-instance. A future v2 could add them as a strict superset of this schema.
  • Terraform has no user-definable named/aliased types. The per-engine include/exclude object shapes are intentionally duplicated character-for-character between source_config.<engine>_source_config and backfill_all.<engine>_excluded_objects — this is by design, not an accidental copy-paste drift risk to "fix" by deduplicating (there is no mechanism to do so in HCL variable type expressions).
  • stream_id/location are force-new, consistent with the house convention that resource-identifying arguments cannot be changed in place on nearly every GCP resource.
  • No IAM propagation concern is unique to this module beyond the general CMEK-grant note above — this module grants no IAM itself; a composition granting a crypto-key role to the Datastream service agent immediately before this module's apply may see a transient permission-denied error for up to ~60 seconds after the grant.

🧱 Design Principles

Concern Secure default Opt-out (explicit)
Stream deletion guard (house extension — no deletion_protection field exists) deletion_policy = "PREVENT" Caller sets "DELETE" or "ABANDON" explicitly
Replication start (desired_state) Left unset → resolves to "NOT_STARTED" — fully configured, but inert Caller sets desired_state = "RUNNING" explicitly
Backfill strategy No default — validation {} requires the caller pick exactly one of backfill_all/backfill_none deliberately N/A — this is a mandatory choice, not an opt-out
CMEK (customer_managed_encryption_key, dataset_template.kms_key_name) Accepted as optional variables — never defaulted to a specific key Caller supplies explicitly, at either or both consumption points
create_without_validation false — Datastream's own live connectivity validation runs at apply Caller sets true explicitly (a real footgun if used to bypass a genuinely broken profile)
Source-engine/connection-profile pairing Not enforceable by this library (plan-only) — the caller is responsible for matching source_config.<engine>_source_config to the actual referenced profile's type N/A — see 🔍 Troubleshooting for the failure mode this can produce

🚀 Runbook

cd C:\GitHubCode\newgooglecloudmodules\terraform-google-datastream-stream
terraform init -backend=false
terraform validate
terraform fmt -check

Pin ?ref=v1.0.0 when consuming this module — never a branch. This library is plan-only; a human applies from CI with valid Workload Identity Federation or ADC credentials.


🧪 Testing

  • terraform init -backend=false, terraform validate, and terraform fmt -check are the entire offline proof gate for this module — all three pass cleanly as of this authoring session, against the actually-resolved hashicorp/google provider version 7.39.0.
  • The single most important limitation of this module's plan-only proof gate — more consequential than the generic "apply-time-only" caveat this library states for every module: validate/fmt CANNOT catch a mismatch between the engine type of the referenced source_connection_profile and this stream's own source_config.<engine>_source_config block. That mismatch passes both cleanly and fails only at apply, surfacing as state = "FAILED" — there is no way for this library, or any plan-only Terraform workflow, to catch it earlier.
  • validate/fmt also cannot catch: an invalid location, a replication_slot/publication that does not actually exist on the source PostgreSQL database, quota/org-policy rejections, or IAM propagation delays on a CMEK grant.
  • The examples/ directory exists so a consuming GitHub Actions workflow can run a real terraform plan against a real project as part of that pipeline's own review gate; this library only guarantees the example is syntactically and structurally sound in isolation.

💬 Example Output

$ terraform output

id = "projects/casey-prod-datastream/locations/us-east1/streams/orders-mysql-to-bq"
state = "RUNNING"
name = "projects/casey-prod-datastream/locations/us-east1/streams/orders-mysql-to-bq"

🔍 Troubleshooting

Symptom Cause Fix
Stream's state output reports "FAILED" even though terraform apply succeeded and desired_state = "RUNNING" The engine of source_config.<engine>_source_config does not match the actual type of the connection profile referenced by source_connection_profile — this passes validate/plan cleanly since Terraform cannot cross-check the two Confirm the referenced connection profile's own profile type (e.g. mysql_profile) matches the engine block set here (e.g. mysql_source_config); fix the mismatch and re-apply
Stream applies successfully but never replicates any data desired_state was left at its default (resolves to "NOT_STARTED") Set desired_state = "RUNNING" explicitly, in this apply or a follow-up one (see Example 8)
Error:... replication slot... does not exist (or similar) at apply on a PostgreSQL source The replication_slot/publication named in postgresql_source_config were never created on the source database — this module and its connection-profile sibling do not create either Create the logical replication slot (with the pgoutput plugin) and the publication on the source PostgreSQL database first, matching the names supplied here, then re-apply
A backfill_all stream is taking far longer than expected, or the source database reports elevated load/quota consumption backfill_all was chosen without considering data volume — it replicates ALL historical data for every included object Consider backfill_none if only forward-looking changes are needed, or narrow include_objects/expand the relevant *_excluded_objects tree to reduce backfill scope
Error: source_config: exactly one of mysql_source_config, oracle_source_config,... at plan time Zero or more than one per-engine source_config block was set Set exactly one of mysql_source_config/oracle_source_config/postgresql_source_config/sql_server_source_config
Error: backfill_all/backfill_none: exactly one must be set... at plan time Both backfill_all and backfill_none were left at their defaults (or both were set) Set exactly one — backfill_all = {... } (possibly {}) XOR backfill_none = true
terraform destroy fails with a deletion_policy error deletion_policy = "PREVENT" (the module default) Expected behavior — change to "ABANDON" or "DELETE" explicitly, re-apply, then destroy
A CMEK-encrypted stream fails to apply with a permission-denied error on the crypto key The Datastream/BigQuery service agent has not been granted a Cloud KMS role on the key, or the grant has not yet propagated (up to ~60 seconds) Confirm the IAM binding exists on the specific crypto key and retry after the propagation window

🔗 Related Docs

  • google_datastream_stream provider reference
  • Primary sibling module: terraform-google-datastream-connection-profile (source + destination profiles this module always requires two instances of)
  • Other siblings: terraform-google-bigquery-dataset, terraform-google-kms-keyring
  • This module's SCOPE.md

💙 "Infrastructure as Code should be standardized, consistent, and secure."

About

Terraform module: terraform-google-datastream-stream

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages