Data Lakehouse

Data store for AI-curated datasets

Purpose & Use Case

The AmSC Data Lakehouse is columnar data storage supporting highly scalable SQL-like data querying.
It is a hosting solution for AI-ready curated datasets generated from raw data.

Functionality

  • Based on Apache Iceberg tables deployed on AWS with underlying Parquet files stored on S3
  • AWS Glue used as lakehouse catalog and integrated with OpenMetadata via Glue connectors
  • PyIceberg can be used to write to Data Lakehouse
  • AWS Data Processing MCP server supports agentic data querying of the lakehouse
  • (tbd) Support for multi-modal data

Current Status

Service available to AmSC Early Users only

Roadmap

MVP ยท Oct 1, 2026

Data catalog search using Lakehouse queries

Data catalog search using Lakehouse queries