TFDS pipeline (Deprecated)

TFDS pipeline (Deprecated)#

Warning

The TFDS input pipeline (dataset_type=tfds) is deprecated. We recommend migrating to the Grain pipeline with dataset_type=grain and grain_file_type=tfrecord. You can keep the same TFRecord dataset paths.

Note

TensorFlow and TensorFlow Datasets (TFDS) are optional dependencies in MaxText. If you need to use the legacy TFDS pipeline, install the optional dependencies by running:

install_tpu_pre_train_extra_deps --with-tf
# or for GPU:
# install_cuda12_pre_train_extra_deps --with-tf
  1. Download the Allenai C4 dataset in TFRecord format to a Cloud Storage bucket. For information about cost, see this discussion

bash download_dataset.sh {GCS_PROJECT} {GCS_BUCKET_NAME}
  1. In src/maxtext/configs/base.yml or through command line, set the following parameters:

dataset_type: tfds
dataset_name: 'c4/en:3.0.1'
# set eval_interval > 0 to use the specified eval dataset. Otherwise, only metrics on the train set will be calculated.
eval_interval: 10000
eval_dataset_name: 'c4/en:3.0.1'
eval_split: 'validation'
# TFDS input pipeline only supports tokenizer in spm format
tokenizer_path: 'src/maxtext/assets/tokenizers/tokenizer.llama2'