> ## Documentation Index
> Fetch the complete documentation index at: https://private-7c7dfe99-sync-clickhouse-operator-docs-7e82242.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

> Documentation for the Parquet format

# Parquet

| Input | Output | Alias |
| ----- | ------ | ----- |
| ✔     | ✔      |       |

<h2 id="description">
  Description
</h2>

[Apache Parquet](https://parquet.apache.org/) is a columnar storage format widespread in the Hadoop ecosystem. ClickHouse supports read and write operations for this format.

<h2 id="data-types-matching-parquet">
  Data types matching
</h2>

The table below shows how Parquet data types match ClickHouse [data types](/core/reference/data-types).

| Parquet type (logical, converted, or physical) | ClickHouse data type                                                                   |
| ---------------------------------------------- | -------------------------------------------------------------------------------------- |
| `BOOLEAN`                                      | [Bool](/core/reference/data-types/boolean)                                             |
| `UINT_8`                                       | [UInt8](/core/reference/data-types/int-uint)                                           |
| `INT_8`                                        | [Int8](/core/reference/data-types/int-uint)                                            |
| `UINT_16`                                      | [UInt16](/core/reference/data-types/int-uint)                                          |
| `INT_16`                                       | [Int16](/core/reference/data-types/int-uint)/[Enum16](/core/reference/data-types/enum) |
| `UINT_32`                                      | [UInt32](/core/reference/data-types/int-uint)                                          |
| `INT_32`                                       | [Int32](/core/reference/data-types/int-uint)                                           |
| `UINT_64`                                      | [UInt64](/core/reference/data-types/int-uint)                                          |
| `INT_64`                                       | [Int64](/core/reference/data-types/int-uint)                                           |
| `DATE`                                         | [Date32](/core/reference/data-types/date)                                              |
| `TIMESTAMP`, `TIME`                            | [DateTime64](/core/reference/data-types/datetime64)                                    |
| `FLOAT`                                        | [Float32](/core/reference/data-types/float)                                            |
| `DOUBLE`                                       | [Float64](/core/reference/data-types/float)                                            |
| `INT96`                                        | [DateTime64(9, 'UTC')](/core/reference/data-types/datetime64)                          |
| `BYTE_ARRAY`, `UTF8`, `ENUM`, `BSON`           | [String](/core/reference/data-types/string)                                            |
| `JSON`                                         | [JSON](/core/reference/data-types/newjson)                                             |
| `FIXED_LEN_BYTE_ARRAY`                         | [FixedString](/core/reference/data-types/fixedstring)                                  |
| `DECIMAL`                                      | [Decimal](/core/reference/data-types/decimal)                                          |
| `LIST`                                         | [Array](/core/reference/data-types/array)                                              |
| `MAP`                                          | [Map](/core/reference/data-types/map)                                                  |
| struct                                         | [Tuple](/core/reference/data-types/tuple)                                              |
| `FLOAT16`                                      | [Float32](/core/reference/data-types/float)                                            |
| `UUID`                                         | [FixedString(16)](/core/reference/data-types/fixedstring)                              |
| `INTERVAL`                                     | [FixedString(12)](/core/reference/data-types/fixedstring)                              |

When writing Parquet file, data types that don't have a matching Parquet type are converted to the nearest available type:

| ClickHouse data type                                                 | Parquet type                                        |
| -------------------------------------------------------------------- | --------------------------------------------------- |
| [IPv4](/core/reference/data-types/ipv4)                              | `UINT_32`                                           |
| [IPv6](/core/reference/data-types/ipv6)                              | `FIXED_LEN_BYTE_ARRAY` (16 bytes)                   |
| [Date](/core/reference/data-types/date) (16 bits)                    | `DATE` (32 bits)                                    |
| [DateTime](/core/reference/data-types/datetime) (32 bits, seconds)   | `TIMESTAMP` (64 bits, milliseconds)                 |
| [Int128/UInt128/Int256/UInt256](/core/reference/data-types/int-uint) | `FIXED_LEN_BYTE_ARRAY` (16/32 bytes, little-endian) |

Arrays can be nested and can have a value of `Nullable` type as an argument. `Tuple` and `Map` types can also be nested.

Data types of ClickHouse table columns can differ from the corresponding fields of the Parquet data inserted. When inserting data, ClickHouse interprets data types according to the table above and then [casts](/core/reference/functions/regular-functions/type-conversion-functions#CAST) the data to that data type which is set for the ClickHouse table column. E.g. a `UINT_32` Parquet column can be read into an [IPv4](/core/reference/data-types/ipv4) ClickHouse column.

For some Parquet types there's no closely matching ClickHouse type. We read them as follows:

* `TIME` (time of day) is read as a timestamp. E.g. `10:23:13.000` becomes `1970-01-01 10:23:13.000`.
* `TIMESTAMP`/`TIME` with `isAdjustedToUTC=false` is a local wall-clock time (year, month, day, hour, minute, second and subsecond fields in a local timezone, regardless of what specific time zone is considered local), same as SQL `TIMESTAMP WITHOUT TIME ZONE`. ClickHouse reads it as if it were a UTC timestamp instead. E.g. `2025-09-29 18:42:13.000` (representing a reading of a local wall clock) becomes `2025-09-29 18:42:13.000` (`DateTime64(3, 'UTC')` representing a point in time). If converted to String, it shows the correct year, month, day, hour, minute, second and subsecond, which can then be interpreted as being in some local timezone instead of UTC. Counterintuitively, changing the type from `DateTime64(3, 'UTC')` to `DateTime64(3)` would not help as both types represent a point in time rather than a clock reading, but `DateTime64(3)` would incorrectly be formatted using local timezone.
* `INTERVAL` is currently read as `FixedString(12)` with raw binary representation of the time interval, as encoded in Parquet file.

<h2 id="example-usage">
  Example usage
</h2>

<h3 id="inserting-data">
  Inserting data
</h3>

Using a Parquet file with the following data, named as `football.parquet`:

```text theme={null}
    ┌───────date─┬─season─┬─home_team─────────────┬─away_team───────────┬─home_team_goals─┬─away_team_goals─┐
 1. │ 2022-04-30 │   2021 │ Sutton United         │ Bradford City       │               1 │               4 │
 2. │ 2022-04-30 │   2021 │ Swindon Town          │ Barrow              │               2 │               1 │
 3. │ 2022-04-30 │   2021 │ Tranmere Rovers       │ Oldham Athletic     │               2 │               0 │
 4. │ 2022-05-02 │   2021 │ Port Vale             │ Newport County      │               1 │               2 │
 5. │ 2022-05-02 │   2021 │ Salford City          │ Mansfield Town      │               2 │               2 │
 6. │ 2022-05-07 │   2021 │ Barrow                │ Northampton Town    │               1 │               3 │
 7. │ 2022-05-07 │   2021 │ Bradford City         │ Carlisle United     │               2 │               0 │
 8. │ 2022-05-07 │   2021 │ Bristol Rovers        │ Scunthorpe United   │               7 │               0 │
 9. │ 2022-05-07 │   2021 │ Exeter City           │ Port Vale           │               0 │               1 │
10. │ 2022-05-07 │   2021 │ Harrogate Town A.F.C. │ Sutton United       │               0 │               2 │
11. │ 2022-05-07 │   2021 │ Hartlepool United     │ Colchester United   │               0 │               2 │
12. │ 2022-05-07 │   2021 │ Leyton Orient         │ Tranmere Rovers     │               0 │               1 │
13. │ 2022-05-07 │   2021 │ Mansfield Town        │ Forest Green Rovers │               2 │               2 │
14. │ 2022-05-07 │   2021 │ Newport County        │ Rochdale            │               0 │               2 │
15. │ 2022-05-07 │   2021 │ Oldham Athletic       │ Crawley Town        │               3 │               3 │
16. │ 2022-05-07 │   2021 │ Stevenage Borough     │ Salford City        │               4 │               2 │
17. │ 2022-05-07 │   2021 │ Walsall               │ Swindon Town        │               0 │               3 │
    └────────────┴────────┴───────────────────────┴─────────────────────┴─────────────────┴─────────────────┘
```

Insert the data:

```sql theme={null}
INSERT INTO football FROM INFILE 'football.parquet' FORMAT Parquet;
```

<h3 id="reading-data">
  Reading data
</h3>

Read data using the `Parquet` format:

```sql theme={null}
SELECT *
FROM football
INTO OUTFILE 'football.parquet'
FORMAT Parquet
```

<Tip>
  Parquet is a binary format that does not display in a human-readable form on the terminal. Use the `INTO OUTFILE` to output Parquet files.
</Tip>

To exchange data with Hadoop, you can use the [`HDFS table engine`](/core/reference/engines/table-engines/integrations/hdfs).

<h2 id="format-settings">
  Format settings
</h2>

| Setting                                                                        | Description                                                                                                                               | Default                                                                                                                                                                                                                                                                                                                                                                                                                                        |
| ------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `input_format_parquet_case_insensitive_column_matching`                        | Ignore case when matching Parquet columns with CH columns.                                                                                | `0`                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `input_format_parquet_preserve_order`                                          | Avoid reordering rows when reading from Parquet files. Usually makes it much slower.                                                      | `0`                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `input_format_parquet_filter_push_down`                                        | When reading Parquet files, skip whole row groups based on the WHERE/PREWHERE expressions and min/max statistics in the Parquet metadata. | `1`                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `input_format_parquet_bloom_filter_push_down`                                  | When reading Parquet files, skip whole row groups based on the WHERE expressions and bloom filter in the Parquet metadata.                | `0`                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `input_format_parquet_allow_missing_columns`                                   | Allow missing columns while reading Parquet input formats                                                                                 | `1`                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `input_format_parquet_local_file_min_bytes_for_seek`                           | Min bytes required for local read (file) to do seek, instead of read with ignore in Parquet input format                                  | `8192`                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| `input_format_parquet_enable_row_group_prefetch`                               | Enable row group prefetching during parquet parsing. Currently, only single-threaded parsing can prefetch.                                | `1`                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `input_format_parquet_skip_columns_with_unsupported_types_in_schema_inference` | Skip columns with unsupported types while schema inference for format Parquet                                                             | `0`                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `input_format_parquet_max_block_size`                                          | Max block size for parquet reader.                                                                                                        | `65409`                                                                                                                                                                                                                                                                                                                                                                                                                                        |
| `input_format_parquet_prefer_block_bytes`                                      | Average block bytes output by parquet reader                                                                                              | `16744704`                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| `input_format_parquet_enable_json_parsing`                                     | When reading Parquet files, parse JSON columns as ClickHouse JSON Column.                                                                 | `1`                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `output_format_parquet_row_group_size`                                         | Target row group size in rows.                                                                                                            | `1000000`                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `output_format_parquet_row_group_size_bytes`                                   | Target row group size in bytes, before compression.                                                                                       | `536870912`                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| `output_format_parquet_string_as_string`                                       | Use Parquet String type instead of Binary for String columns.                                                                             | `1`                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `output_format_parquet_fixed_string_as_fixed_byte_array`                       | Use Parquet FIXED\_LEN\_BYTE\_ARRAY type instead of Binary for FixedString columns.                                                       | `1`                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `output_format_parquet_compression_method`                                     | Compression method for Parquet output format. Supported codecs: snappy, lz4, brotli, zstd, gzip, none (uncompressed)                      | `zstd`                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| `output_format_parquet_parallel_encoding`                                      | Do Parquet encoding in multiple threads.                                                                                                  | `1`                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `output_format_parquet_data_page_size`                                         | Target page size in bytes, before compression.                                                                                            | `1048576`                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| `output_format_parquet_batch_size`                                             | Check page size every this many rows. Consider decreasing if you have columns with average values size above a few KBs.                   | `1024`                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| `output_format_parquet_write_page_index`                                       | Add a possibility to write page index into parquet files.                                                                                 | `1`                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `input_format_parquet_import_nested`                                           | Obsolete setting, does nothing.                                                                                                           | `0`                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| `input_format_parquet_local_time_as_utc`                                       | true                                                                                                                                      | Determines the data type used by schema inference for Parquet timestamps with isAdjustedToUTC=false. If true: DateTime64(..., 'UTC'), if false: DateTime64(...). Neither behavior is fully correct as ClickHouse doesn't have a data type for local wall-clock time. Counterintuitively, 'true' is probably the less incorrect option, because formatting the 'UTC' timestamp as String will produce representation of the correct local time. |
