Purpose & Use Case
The AmSC Data Lakehouse is columnar data storage supporting highly scalable SQL-like data querying.
It is a hosting solution for AI-ready curated datasets generated from raw data.
Functionality
- Based on Apache Iceberg tables deployed on AWS with underlying Parquet files stored on S3
- AWS Glue used as lakehouse catalog and integrated with OpenMetadata via Glue connectors
- PyIceberg can be used to write to Data Lakehouse
- AWS Data Processing MCP server supports agentic data querying of the lakehouse
- (tbd) Support for multi-modal data
Current Status
Service available to AmSC Early Users only
Roadmap
MVP ยท Oct 1, 2026
Data catalog search using Lakehouse queries
Data catalog search using Lakehouse queries