Skip to main content

to_export

Recommended: Universal export function supporting multiple output formats with schema validation and customization options.

Installation

Basic Usage

Parameters

  • data (pd.DataFrame): DataFrame to export
  • name (str): Output file name/stream name
  • output_dir (str): Directory for output files
  • keys (list): Primary key fields
  • unified_model (pydantic.BaseModel): Pydantic model for schema validation
  • export_format (str): Output format (‘singer’, ‘parquet’, ‘json’, ‘jsonl’, ‘csv’)
  • output_file_prefix (str): Optional prefix for output files
  • schema (dict): Custom schema for Singer format
  • stringify_objects (bool): Convert complex objects to strings for Parquet
  • target_state_fields (list | str): Column names from export records to persist in target state customData (Singer export only)
  • target_state_include_hash (bool): Persist the target SDK record hash in target state customData (Singer export only)

Returns

None (writes files to specified directory)

Notes

  • Uses environment variables for format defaults
  • Supports prefix override per stream
  • Handles complex data types appropriately per format
  • Validates against Pydantic models when provided

to_singer

Recommended: Specialized function for exporting data to Singer format with comprehensive type handling.

Usage

Parameters

  • df (pd.DataFrame): DataFrame to export
  • stream (str): Singer stream name
  • output_dir (str): Output directory
  • keys (list): Primary key fields
  • filename (str): Output filename (default: ‘data.singer’)
  • allow_objects (bool): Enable complex object handling
  • schema (dict): Custom schema definition
  • unified_model (pydantic.BaseModel): Pydantic model for validation
  • target_state_fields (list | str): Column names from export records to persist in target state customData
  • target_state_include_hash (bool): Persist the target SDK record hash in customData

Notes

  • Handles datetime conversions automatically
  • Supports catalog-based schemas
  • Validates data against provided schemas
  • Creates Singer-compliant output files

Target state metadata

Pass target_state_fields and/or target_state_include_hash when you want hotglue to store extra per-record data on the target state after a write job. The same options work on to_singer and to_export when export_format is singer.
  • target_state_fields: columns to persist in target state customData. Field names in target_state_fields refer to columns on records as your ETL sends them to the target. A single string is treated as one field.
  • target_state_include_hash: persist the target SDK record hash in target state customData.
When you set these options, Gluestick adds an x-hotglue block on the stream SCHEMA message in data.singer. The target SDK reads that configuration and copies the requested values into each record’s customData in the target state output (merged with any customData the target already emits). That customData is also stored in the tenant snapshots folder when Save records IDs in snapshots is enabled on the flow. The usual reason to set this on export is so a later transformation run can read what was stored on the last write job. Typical uses include comparing the saved hash (target_state_include_hash) to detect unchanged records, reusing values you pushed previously, or joining snapshot fields to incoming sync data on externalId before the next export. Example: load the write snapshot and merge CustomData fields that match your target_state_fields, plus hash when you enabled target_state_include_hash:

Common Patterns

Handling Complex Exports