remove torch install for now

address pr feedback
gptq doesn't play well with sample packing
2023-09-01 08:15:22 -07:00 · 2023-08-29 22:45:22 -07:00 · 2023-08-29 12:15:31 -07:00 · 2023-08-29 12:02:51 -07:00 · 2023-08-29 12:02:51 -07:00 · 2023-08-29 12:02:50 -07:00
34 changed files with 252 additions and 1011 deletions
--- a/.github/workflows/pypi.yml
+++ b/.github/workflows/pypi.yml
@@ -1,45 +0,0 @@
-name: publish pypi
-
-on:
-  push:
-    tags:
-      - '*'
-
-jobs:
-  pypi-publish:
-    name: Upload release to PyPI
-    runs-on: ubuntu-latest
-    environment:
-      name: pypi
-      url: https://pypi.org/p/axolotl
-    permissions:
-      id-token: write  # IMPORTANT: this permission is mandatory for trusted publishing
-    steps:
-      - name: Check out repository code
-        uses: actions/checkout@v3
-
-      - name: Setup Python
-        uses: actions/setup-python@v4
-        with:
-          python-version: "3.10"
-
-      - name: Install dependencies
-        run: |
-          pip3 install wheel
-          pip3 install -e .
-          pip3 install -r requirements-tests.txt
-
-      - name: Extract tag name
-        id: tag
-        run: echo ::set-output name=TAG_NAME::$(echo $GITHUB_REF | cut -d / -f 3)
-
-      - name: Update version in setup.py
-        run: >-
-          sed -i -E 's/version="([0-9.]+)",/version="${{ steps.tag.outputs.TAG_NAME }}",/g' setup.py
-
-      - name: Build a binary wheel
-        run: >-
-          python setup.py sdist bdist_wheel
-
-      - name: Publish package distributions to PyPI
-        uses: pypa/gh-action-pypi-publish@release/v1
--- a/.github/workflows/tests.yml
+++ b/.github/workflows/tests.yml
@@ -24,8 +24,8 @@ jobs:

      - name: Install dependencies
        run: |
-          pip3 install -e .
-          pip3 install -r requirements-tests.txt
+          pip install -e .
+          pip install -r requirements-tests.txt

      - name: Run tests
        run: |
--- a/README.md
+++ b/README.md
@@ -90,7 +90,8 @@ accelerate launch scripts/finetune.py examples/openllama-3b/lora.yml \
  ```bash
  docker run --gpus '"all"' --rm -it winglian/axolotl:main-py3.10-cu118-2.0.1
  ```
-  - `winglian/axolotl-runpod:main-latest`: for runpod or use this [direct link](https://runpod.io/gsc?template=v2ickqhz9s&ref=6i7fkpdz)
+  - `winglian/axolotl-runpod:main-py3.10-cu118-2.0.1`: for runpod
+  - `winglian/axolotl-runpod:main-py3.9-cu118-2.0.1-gptq`: for gptq

  Or run on the current files for development:

@@ -103,9 +104,19 @@ accelerate launch scripts/finetune.py examples/openllama-3b/lora.yml \

  2. Install pytorch stable https://pytorch.org/get-started/locally/

-  3. Install axolotl along with python dependencies
+  3. Install python dependencies with ONE of the following:
+      - Recommended, supports QLoRA, NO gptq/int4 support
        ```bash
-        pip3 install -e .[flash-attn]
+        pip3 install -e .
+        pip3 install -U git+https://github.com/huggingface/peft.git
+        ```
+      - gptq/int4 support, NO QLoRA
+        ```bash
+        pip3 install -e .[gptq]
+        ```
+      - same as above but not recommended
+        ```bash
+        pip3 install -e .[gptq_triton]
        ```

 - LambdaLabs
@@ -140,9 +151,10 @@ accelerate launch scripts/finetune.py examples/openllama-3b/lora.yml \
  git clone https://github.com/OpenAccess-AI-Collective/axolotl
  cd axolotl

-  pip3 install -e .
+  pip3 install -e . # change depend on needs
  pip3 install protobuf==3.20.3
  pip3 install -U --ignore-installed requests Pillow psutil scipy
+  pip3 install git+https://github.com/huggingface/peft.git # not for gptq
  ```

  5. Set path
@@ -151,8 +163,6 @@ accelerate launch scripts/finetune.py examples/openllama-3b/lora.yml \
  ```
  </details>

- Windows: Please use WSL or Docker!
-
 ### Dataset

 Axolotl supports a variety of dataset formats. Below are some of the formats you can use.
@@ -560,30 +570,6 @@ log_sweep_min_lr:
 log_sweep_max_lr:

 # specify optimizer
-# Valid values are driven by the Transformers OptimizerNames class, see:
-# https://github.com/huggingface/transformers/blob/95b374952dc27d8511541d6f5a4e22c9ec11fb24/src/transformers/training_args.py#L134
-#
-# Note that not all optimizers may be available in your environment, ex: 'adamw_anyprecision' is part of
-# torchdistx, 'adamw_bnb_8bit' is part of bnb.optim.Adam8bit, etc. When in doubt, it is recommended to start with the optimizer used
-# in the examples/ for your model and fine-tuning use case.
-#
-# Valid values for 'optimizer' include:
-# - adamw_hf
-# - adamw_torch
-# - adamw_torch_fused
-# - adamw_torch_xla
-# - adamw_apex_fused
-# - adafactor
-# - adamw_anyprecision
-# - sgd
-# - adagrad
-# - adamw_bnb_8bit
-# - lion_8bit
-# - lion_32bit
-# - paged_adamw_32bit
-# - paged_adamw_8bit
-# - paged_lion_32bit
-# - paged_lion_8bit
 optimizer:
 # specify weight decay
 weight_decay:
@@ -637,11 +623,6 @@ fsdp_config:
 # Deepspeed config path
 deepspeed:

-# Advanced DDP Arguments
-ddp_timeout:
-ddp_bucket_cap_mb:
-ddp_broadcast_buffers:
-
 # Path to torch distx for optim 'adamw_anyprecision'
 torchdistx_path:

@@ -764,10 +745,6 @@ Try to turn off xformers.

 It's safe to ignore it.

-> NCCL Timeouts during training
-
-See the [NCCL](docs/nccl.md) guide.
-
 ## Need help? 🙋♂️

 Join our [Discord server](https://discord.gg/HhrNrHJPRb) where we can help you
--- a/deepspeed/zero3.json
+++ b/deepspeed/zero3.json
@@ -35,7 +35,10 @@
    "type": "AdamW",
    "params": {
      "lr": "auto",
-      "betas": "auto",
+      "betas": [
+        0.9,
+        0.95
+      ],
      "eps": 1e-8,
      "weight_decay": "auto"
    }
--- a/docker-compose.yaml
+++ b/docker-compose.yaml
@@ -9,11 +9,6 @@ services:
      - ~/.cache/huggingface/:/root/.cache/huggingface/
    # set environment variables
    environment:
-      # Set environment variables
-      - GIT_AUTHOR_NAME=${GIT_AUTHOR_NAME}
-      - GIT_AUTHOR_EMAIL=${GIT_AUTHOR_EMAIL}
-      - GIT_COMMITTER_NAME=${GIT_COMMITTER_NAME}
-      - GIT_COMMITTER_EMAIL=${GIT_COMMITTER_EMAIL}
      - WANDB_API_KEY=${WANDB_API_KEY}
    deploy:
      resources:
--- a/docker/Dockerfile
+++ b/docker/Dockerfile
@@ -15,9 +15,9 @@ RUN git clone --depth=1 https://github.com/OpenAccess-AI-Collective/axolotl.git
 # If AXOLOTL_EXTRAS is set, append it in brackets
 RUN cd axolotl && \
    if [ "$AXOLOTL_EXTRAS" != "" ] ; then \
-        pip install -e .[flash-attn,$AXOLOTL_EXTRAS]; \
+        pip install -e .[flash-attn,gptq,$AXOLOTL_EXTRAS]; \
    else \
-        pip install -e .[flash-attn]; \
+        pip install -e .[flash-attn,gptq]; \
    fi

 # fix so that git fetch/pull from remote works
--- a/docs/nccl.md
+++ b/docs/nccl.md
@@ -1,46 +0,0 @@
-# NCCL
-
-NVIDIA NCCL is a library to facilitate and optimize multi-GPU communication operations, such as broadcast, all-gather, reduce, all-reduce, etc. Broadly, NCCL configuration is highly environment-specific and is configured via several [environment variables](https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/env.html). A common NCCL-related problem occurs when a long-running operation times out causing the training process to abort:
-
-```text
-Watchdog caught collective operation timeout: WorkNCCL(SeqNum=42, OpType=ALLGATHER, Timeout(ms)=1800000) ran for 1806948 milliseconds before timing out.
-```
-
-Often, this timeout will happen after 30 minutes (the default setting) and is accompanied by below-average power consumption with near 100% GPU utilization before the error is raised. Nvidia recommends [disabling PCI access control services (ACS)](https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/troubleshooting.html#pci-access-control-services-acs) as a possible solution if this is available to you.
-
-Forcing cross-GPU communication via [NVLink](https://en.wikipedia.org/wiki/NVLink) may help without increasing timeouts. To verify that your configuration is leveraging NVLink run the following command:
-
-```shell
-nvidia-smi nvlink --status
-```
-
-To force NCCL to use NVLink, simply set this in the environment:
-
-```shell
-export NCCL_P2P_LEVEL=NVL
-```
-
-If NVLink is not available in your environment there are other options for ``NCCL_P2P_LEVEL`` in the table below:
-
-| NCCL_P2P_LEVEL | Description |
-| -------------- | ----------- |
-| PIX | P2P data transfers through no more than a single PCIe bridge. Faster data transfer rates vs to paths involving multiple bridges, but slower compared to direct GPU-to-GPU communication. |
-| PXB | P2P data transfers through multiple PCIe bridges but not going through the PCIe Host Bridge; this path involves a complex routing process, potentially incurring a moderate level of latency. |
-| PHB | P2P data transfers occur over the PCIe and through a PCIe Host Bridge, typically involving the CPU, which can facilitate direct memory access but might introduce additional latency compared to more direct paths (ex PIX, NVL) |
-
-To validate that acceptable data transfer speeds exist for your training job, running [NCCL Tests](https://github.com/NVIDIA/nccl-tests/blob/master/README.md) can help pinpoint bottlenecks, for example:
-
-```shell
-./build/all_reduce_perf -b 8 -e 128M -f 2 -g 3
-```
-
-It can be useful when debugging NCCL communication timeouts to activate additional logging in both PyTorch and NCCL:
-
-```shell
-export NCCL_DEBUG=INFO
-export NCCL_DEBUG_SUBSYS=ALL
-export TORCH_DISTRIBUTED_DEBUG=INFO
-export TORCHELASTIC_ERROR_FILE=/PATH/TO/torcherror.log
-```
-
-Finally, if you believe your training job needs more time you can increase the timeout past 30 minutes by setting the ``ddp_timeout`` value in the Axolotl configuration. See [PyTorch init_process_group](https://pytorch.org/docs/stable/distributed.html#torch.distributed.init_process_group) for documentation on this value.
--- a/examples/code-llama/13b/lora.yml
+++ b/examples/code-llama/13b/lora.yml
@@ -17,7 +17,6 @@ output_dir: ./lora-out

 sequence_len: 100000
 sample_packing: true
-pad_to_sequence_len: true

 adapter: lora
 lora_model_dir:
--- a/examples/code-llama/13b/qlora.yml
+++ b/examples/code-llama/13b/qlora.yml
@@ -20,7 +20,6 @@ lora_model_dir:

 sequence_len: 100000
 sample_packing: true
-pad_to_sequence_len: true

 lora_r: 32
 lora_alpha: 16
--- a/examples/code-llama/34b/lora.yml
+++ b/examples/code-llama/34b/lora.yml
@@ -17,7 +17,6 @@ output_dir: ./lora-out

 sequence_len: 100000
 sample_packing: true
-pad_to_sequence_len: true

 adapter: lora
 lora_model_dir:
--- a/examples/code-llama/34b/qlora.yml
+++ b/examples/code-llama/34b/qlora.yml
@@ -20,7 +20,6 @@ lora_model_dir:

 sequence_len: 100000
 sample_packing: true
-pad_to_sequence_len: true

 lora_r: 32
 lora_alpha: 16
--- a/examples/code-llama/7b/lora.yml
+++ b/examples/code-llama/7b/lora.yml
@@ -17,7 +17,6 @@ output_dir: ./lora-out

 sequence_len: 100000
 sample_packing: true
-pad_to_sequence_len: true

 adapter: lora
 lora_model_dir:
--- a/examples/code-llama/7b/qlora.yml
+++ b/examples/code-llama/7b/qlora.yml
@@ -20,7 +20,6 @@ lora_model_dir:

 sequence_len: 100000
 sample_packing: true
-pad_to_sequence_len: true

 lora_r: 32
 lora_alpha: 16
--- a/examples/llama-2/lora.yml
+++ b/examples/llama-2/lora.yml
@@ -17,7 +17,6 @@ output_dir: ./lora-out

 sequence_len: 4096
 sample_packing: true
-pad_to_sequence_len: true

 adapter: lora
 lora_model_dir:
--- a/examples/llama-2/qlora.yml
+++ b/examples/llama-2/qlora.yml
@@ -20,7 +20,6 @@ lora_model_dir:

 sequence_len: 4096
 sample_packing: true
-pad_to_sequence_len: true

 lora_r: 32
 lora_alpha: 16
--- a/examples/llama-2/relora.yml
+++ b/examples/llama-2/relora.yml
@@ -20,7 +20,6 @@ lora_model_dir:

 sequence_len: 4096
 sample_packing: true
-pad_to_sequence_len: true

 lora_r: 8
 lora_alpha: 16
--- a/requirements.txt
+++ b/requirements.txt
@@ -6,13 +6,12 @@ packaging
 peft @ git+https://github.com/huggingface/peft.git
 transformers @ git+https://github.com/huggingface/transformers.git
 bitsandbytes>=0.41.1
-accelerate @ git+https://github.com/huggingface/accelerate
+accelerate @ git+https://github.com/huggingface/accelerate@2a289f6108e77a77a4efffb3f6316bc98538413b
 addict
-evaluate
 fire
 PyYAML>=6.0
 datasets
-flash-attn>=2.2.1
+flash-attn>=2.0.8
 sentencepiece
 wandb
 einops
--- a/scripts/finetune.py
+++ b/scripts/finetune.py
@@ -4,7 +4,9 @@ import importlib
 import logging
 import os
 import random
+import signal
 import sys
+from dataclasses import dataclass, field
 from pathlib import Path
 from typing import Any, Dict, List, Optional, Union

@@ -15,17 +17,17 @@ import yaml

 # add src to the pythonpath so we don't need to pip install this
 from art import text2art
+from optimum.bettertransformer import BetterTransformer
 from transformers import GenerationConfig, TextStreamer

-from axolotl.common.cli import TrainerCliArgs, load_model_and_tokenizer
 from axolotl.logging_config import configure_logging
-from axolotl.train import TrainDatasetMeta, train
 from axolotl.utils.config import normalize_config, validate_config
 from axolotl.utils.data import prepare_dataset
 from axolotl.utils.dict import DictDefault
 from axolotl.utils.distributed import is_main_process
-from axolotl.utils.models import load_tokenizer
+from axolotl.utils.models import load_model, load_model_config, load_tokenizer
 from axolotl.utils.tokenization import check_dataset_labels
+from axolotl.utils.trainer import setup_trainer
 from axolotl.utils.wandb import setup_wandb_env_vars

 project_root = os.path.abspath(os.path.join(os.path.dirname(__file__), ".."))
@@ -38,13 +40,26 @@ LOG = logging.getLogger("axolotl.scripts")
 os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"


+@dataclass
+class TrainerCliArgs:
+    """
+    dataclass representing the various non-training arguments
+    """
+
+    debug: bool = field(default=False)
+    inference: bool = field(default=False)
+    merge_lora: bool = field(default=False)
+    prepare_ds_only: bool = field(default=False)
+    prompter: Optional[str] = field(default=None)
+    shard: bool = field(default=False)
+
+
 def print_axolotl_text_art(suffix=None):
    font = "nancyj"
    ascii_text = "  axolotl"
    if suffix:
        ascii_text += f"  x  {suffix}"
    ascii_art = text2art(" axolotl", font=font)
-
    if is_main_process():
        print(ascii_art)

@@ -58,45 +73,9 @@ def get_multi_line_input() -> Optional[str]:
    return instruction


-def do_merge_lora(
-    *,
-    cfg: DictDefault,
-    cli_args: TrainerCliArgs,
-):
-    model, tokenizer = load_model_and_tokenizer(cfg=cfg, cli_args=cli_args)
-    safe_serialization = cfg.save_safetensors is True
-
-    LOG.info("running merge of LoRA with base model")
-    model = model.merge_and_unload()
-    model.to(dtype=torch.float16)
-
-    if cfg.local_rank == 0:
-        LOG.info("saving merged model")
-        model.save_pretrained(
-            str(Path(cfg.output_dir) / "merged"),
-            safe_serialization=safe_serialization,
-        )
-        tokenizer.save_pretrained(str(Path(cfg.output_dir) / "merged"))
-
-
-def shard(
-    *,
-    cfg: DictDefault,
-    cli_args: TrainerCliArgs,
-):
-    model, _ = load_model_and_tokenizer(cfg=cfg, cli_args=cli_args)
-    safe_serialization = cfg.save_safetensors is True
-    LOG.debug("Re-saving model w/ sharding")
-    model.save_pretrained(cfg.output_dir, safe_serialization=safe_serialization)
-
-
-def do_inference(
-    *,
-    cfg: DictDefault,
-    cli_args: TrainerCliArgs,
-):
-    model, tokenizer = load_model_and_tokenizer(cfg=cfg, cli_args=cli_args)
-    prompter = cli_args.prompter
+def do_inference(cfg, model, tokenizer, prompter: Optional[str]):
+    if prompter == "None":
+        prompter = None
    default_tokens = {"unk_token": "<unk>", "bos_token": "<s>", "eos_token": "</s>"}

    for token, symbol in default_tokens.items():
@@ -197,6 +176,141 @@ def check_not_in(list1: List[str], list2: Union[Dict[str, Any], List[str]]) -> b
    return not any(el in list2 for el in list1)


+def train(
+    *,
+    cfg: DictDefault,
+    cli_args: TrainerCliArgs,
+):
+    # load the tokenizer first
+    LOG.info(f"loading tokenizer... {cfg.tokenizer_config or cfg.base_model_config}")
+    tokenizer = load_tokenizer(cfg)
+
+    if not (
+        cli_args.shard or cli_args.merge_lora or cli_args.inference
+    ):  # don't need to load dataset for these
+        train_dataset, eval_dataset, total_num_steps = prepare_dataset(cfg, tokenizer)
+
+    if cli_args.debug or cfg.debug:
+        LOG.info("check_dataset_labels...")
+        check_dataset_labels(
+            train_dataset.select(
+                [random.randrange(0, len(train_dataset) - 1) for _ in range(5)]  # nosec
+            ),
+            tokenizer,
+        )
+
+    if cli_args.prepare_ds_only:
+        LOG.info("Finished preparing dataset. Exiting...")
+        return
+
+    # Load the model and tokenizer
+    LOG.info("loading model and (optionally) peft_config...")
+    model, peft_config = load_model(cfg, tokenizer, inference=cli_args.inference)
+
+    safe_serialization = cfg.save_safetensors is True
+
+    if cli_args.merge_lora and cfg.adapter is not None:
+        LOG.info("running merge of LoRA with base model")
+        model = model.merge_and_unload()
+        model.to(dtype=torch.float16)
+
+        if cfg.local_rank == 0:
+            LOG.info("saving merged model")
+            model.save_pretrained(
+                str(Path(cfg.output_dir) / "merged"),
+                safe_serialization=safe_serialization,
+            )
+            tokenizer.save_pretrained(str(Path(cfg.output_dir) / "merged"))
+        return
+
+    if cli_args.inference:
+        LOG.debug("Running inference on model")
+        do_inference(cfg, model, tokenizer, prompter=cli_args.prompter)
+        return
+
+    if cli_args.shard:
+        LOG.debug("Re-saving model w/ sharding")
+        model.save_pretrained(cfg.output_dir, safe_serialization=safe_serialization)
+        return
+
+    if cfg.resume_from_checkpoint is None and cfg.auto_resume_from_checkpoints:
+        possible_checkpoints = [
+            str(cp) for cp in Path(cfg.output_dir).glob("checkpoint-*")
+        ]
+        if len(possible_checkpoints) > 0:
+            sorted_paths = sorted(
+                possible_checkpoints,
+                key=lambda path: int(path.split("-")[-1]),
+            )
+            cfg.resume_from_checkpoint = sorted_paths[-1]
+            LOG.info(
+                f"Using Auto-resume functionality to start with checkpoint at {cfg.resume_from_checkpoint}"
+            )
+    resume_from_checkpoint = cfg.resume_from_checkpoint
+
+    trainer = setup_trainer(
+        cfg, train_dataset, eval_dataset, model, tokenizer, total_num_steps
+    )
+
+    model.config.use_cache = False
+
+    if torch.__version__ >= "2" and sys.platform != "win32":
+        LOG.info("Compiling torch model")
+        model = torch.compile(model)
+
+    # go ahead and presave, so we have the adapter config available to inspect
+    if peft_config:
+        LOG.info(f"Pre-saving adapter config to {cfg.output_dir}")
+        peft_config.save_pretrained(cfg.output_dir)
+
+    # In case we want to stop early with ctrl+c, this is a nice to have to save the pretrained model
+    if cfg.local_rank == 0:
+
+        def terminate_handler(_, __, model):
+            if cfg.flash_optimum:
+                model = BetterTransformer.reverse(model)
+            model.save_pretrained(cfg.output_dir, safe_serialization=safe_serialization)
+            sys.exit(0)
+
+        signal.signal(
+            signal.SIGINT, lambda signum, frame: terminate_handler(signum, frame, model)
+        )
+
+    LOG.info("Starting trainer...")
+    if cfg.group_by_length:
+        LOG.info("hang tight... sorting dataset for group_by_length")
+
+    if not Path(cfg.output_dir).is_dir():
+        os.makedirs(cfg.output_dir, exist_ok=True)
+    tokenizer.save_pretrained(cfg.output_dir)
+    if cfg.flash_optimum:
+        with torch.backends.cuda.sdp_kernel(
+            enable_flash=True, enable_math=True, enable_mem_efficient=True
+        ):
+            trainer.train(resume_from_checkpoint=resume_from_checkpoint)
+    else:
+        trainer.train(resume_from_checkpoint=resume_from_checkpoint)
+
+    LOG.info(f"Training Completed!!! Saving pre-trained model to {cfg.output_dir}")
+
+    if cfg.relora_steps:
+        if cfg.adapter == "lora" and not (cfg.load_in_4bit or cfg.load_in_8bit):
+            model = model.merge_and_unload()
+        else:
+            # final model weights have already been saved by `ReLoRACallback.on_train_end`
+            return
+
+    # TODO do we need this fix? https://huggingface.co/docs/accelerate/usage_guides/fsdp#saving-and-loading
+    # only save on rank 0, otherwise it corrupts output on multi-GPU when multiple processes attempt to write the same file
+    if cfg.fsdp:
+        trainer.save_model(cfg.output_dir)
+    elif cfg.local_rank == 0:
+        if cfg.flash_optimum:
+            model = BetterTransformer.reverse(model)
+
+        model.save_pretrained(cfg.output_dir, safe_serialization=safe_serialization)
+
+
 def load_cfg(config: Path = Path("examples/"), **kwargs):
    if Path(config).is_dir():
        config = choose_config(config)
@@ -216,6 +330,15 @@ def load_cfg(config: Path = Path("examples/"), **kwargs):
            else:
                cfg[k] = kwargs[k]

+    model_config = load_model_config(cfg)
+
+    # figure out if the model is llama
+    cfg.is_llama_derived_model = (
+        (hasattr(model_config, "model_type") and model_config.model_type == "llama")
+        or cfg.is_llama_derived_model
+        or "llama" in cfg.base_model
+        or (cfg.model_type and "llama" in cfg.model_type.lower())
+    )
    validate_config(cfg)

    normalize_config(cfg)
@@ -224,55 +347,15 @@ def load_cfg(config: Path = Path("examples/"), **kwargs):
    return cfg


-def load_datasets(
-    *,
-    cfg: DictDefault,
-    cli_args: TrainerCliArgs,
-) -> TrainDatasetMeta:
-    tokenizer = load_tokenizer(cfg)
-
-    train_dataset, eval_dataset, total_num_steps = prepare_dataset(cfg, tokenizer)
-
-    if cli_args.debug or cfg.debug:
-        LOG.info("check_dataset_labels...")
-        check_dataset_labels(
-            train_dataset.select(
-                [
-                    random.randrange(0, len(train_dataset) - 1)  # nosec
-                    for _ in range(cli_args.debug_num_examples)
-                ]
-            ),
-            tokenizer,
-            num_examples=cli_args.debug_num_examples,
-            text_only=cli_args.debug_text_only,
-        )
-
-    return TrainDatasetMeta(
-        train_dataset=train_dataset,
-        eval_dataset=eval_dataset,
-        total_num_steps=total_num_steps,
-    )
-
-
-def do_cli(config: Path = Path("examples/"), **kwargs):
+def do_train(config: Path = Path("examples/"), **kwargs):
    print_axolotl_text_art()
    parsed_cfg = load_cfg(config, **kwargs)
    parser = transformers.HfArgumentParser((TrainerCliArgs))
    parsed_cli_args, _ = parser.parse_args_into_dataclasses(
        return_remaining_strings=True
    )
-    if parsed_cli_args.inference:
-        do_inference(cfg=parsed_cfg, cli_args=parsed_cli_args)
-    elif parsed_cli_args.merge_lora:
-        do_merge_lora(cfg=parsed_cfg, cli_args=parsed_cli_args)
-    elif parsed_cli_args.shard:
-        shard(cfg=parsed_cfg, cli_args=parsed_cli_args)
-    else:
-        dataset_meta = load_datasets(cfg=parsed_cfg, cli_args=parsed_cli_args)
-        if parsed_cli_args.prepare_ds_only:
-            return
-        train(cfg=parsed_cfg, cli_args=parsed_cli_args, dataset_meta=dataset_meta)
+    train(cfg=parsed_cfg, cli_args=parsed_cli_args)


 if __name__ == "__main__":
-    fire.Fire(do_cli)
+    fire.Fire(do_train)
--- a/setup.py
+++ b/setup.py
@@ -7,7 +7,9 @@ def parse_requirements():
    _install_requires = []
    _dependency_links = []
    with open("./requirements.txt", encoding="utf-8") as requirements_file:
-        lines = [r.strip() for r in requirements_file.readlines()]
+        lines = [
+            r.strip() for r in requirements_file.readlines() if "auto-gptq" not in r
+        ]
        for line in lines:
            if line.startswith("--extra-index-url"):
                # Handle custom index URLs
@@ -24,16 +26,18 @@ install_requires, dependency_links = parse_requirements()

 setup(
    name="axolotl",
-    version="0.3.0",
-    description="LLM Trainer",
-    long_description="Axolotl is a tool designed to streamline the fine-tuning of various AI models, offering support for multiple configurations and architectures.",
+    version="0.1",
+    description="You know you're going to axolotl questions",
    package_dir={"": "src"},
    packages=find_packages(),
    install_requires=install_requires,
    dependency_links=dependency_links,
    extras_require={
+        "gptq": [
+            "auto-gptq",
+        ],
        "flash-attn": [
-            "flash-attn>=2.2.1",
+            "flash-attn==2.0.8",
        ],
        "extras": [
            "deepspeed",
--- a/src/axolotl/common/init.py
+++ b/src/axolotl/common/init.py
--- a/src/axolotl/common/cli.py
+++ b/src/axolotl/common/cli.py
@@ -1,43 +0,0 @@
-"""
-shared module for cli specific things
-"""
-
-import logging
-from dataclasses import dataclass, field
-from typing import Optional
-
-from axolotl.logging_config import configure_logging
-from axolotl.utils.dict import DictDefault
-from axolotl.utils.models import load_model, load_tokenizer
-
-configure_logging()
-LOG = logging.getLogger("axolotl.common.cli")
-
-
-@dataclass
-class TrainerCliArgs:
-    """
-    dataclass representing the various non-training arguments
-    """
-
-    debug: bool = field(default=False)
-    debug_text_only: bool = field(default=False)
-    debug_num_examples: int = field(default=5)
-    inference: bool = field(default=False)
-    merge_lora: bool = field(default=False)
-    prepare_ds_only: bool = field(default=False)
-    prompter: Optional[str] = field(default=None)
-    shard: bool = field(default=False)
-
-
-def load_model_and_tokenizer(
-    *,
-    cfg: DictDefault,
-    cli_args: TrainerCliArgs,
-):
-    LOG.info(f"loading tokenizer... {cfg.tokenizer_config or cfg.base_model_config}")
-    tokenizer = load_tokenizer(cfg)
-    LOG.info("loading model and (optionally) peft_config...")
-    model, _ = load_model(cfg, tokenizer, inference=cli_args.inference)
-
-    return model, tokenizer
--- a/src/axolotl/logging_config.py
+++ b/src/axolotl/logging_config.py
@@ -23,7 +23,6 @@ class ColorfulFormatter(Formatter):
    }

    def format(self, record):
-        record.rank = int(os.getenv("LOCAL_RANK", "0"))
        log_message = super().format(record)
        return self.COLORS.get(record.levelname, "") + log_message + Fore.RESET

@@ -36,7 +35,7 @@ DEFAULT_LOGGING_CONFIG: Dict[str, Any] = {
        },
        "colorful": {
            "()": ColorfulFormatter,
-            "format": "[%(asctime)s] [%(levelname)s] [%(name)s.%(funcName)s:%(lineno)d] [PID:%(process)d] [RANK:%(rank)d] %(message)s",
+            "format": "[%(asctime)s] [%(levelname)s] [%(name)s.%(funcName)s:%(lineno)d] [PID:%(process)d] %(message)s",
        },
    },
    "filters": {},
--- a/src/axolotl/monkeypatch/llama_attn_hijack_flash.py
+++ b/src/axolotl/monkeypatch/llama_attn_hijack_flash.py
@@ -2,9 +2,7 @@

 # copied from https://github.com/lm-sys/FastChat/blob/main/fastchat/train/llama_flash_attn_monkey_patch.py

-import logging
 import warnings
-from functools import partial
 from typing import List, Optional, Tuple, Union

 import torch
@@ -35,9 +33,6 @@ except ImportError:
    )


-LOG = logging.getLogger("axolotl")
-
-
 def replace_llama_attn_with_flash_attn(packed: Optional[bool] = False):
    transformers.models.llama.modeling_llama.LlamaModel._prepare_decoder_attention_mask = (  # pylint: disable=protected-access
        _prepare_decoder_attention_mask
@@ -49,34 +44,6 @@ def replace_llama_attn_with_flash_attn(packed: Optional[bool] = False):
            llama_model_forward
        )

-    try:
-        from flash_attn.losses.cross_entropy import CrossEntropyLoss
-
-        LOG.info("patching with flash_attn.losses.cross_entropy")
-        transformers.models.llama.modeling_llama.CrossEntropyLoss = partial(
-            CrossEntropyLoss, inplace_backward=True
-        )
-    except ImportError:
-        LOG.info(
-            "optimized flash-attention CrossEntropyLoss not found (run `pip install 'git+https://github.com/Dao-AILab/flash-attention.git#egg=xentropy_cuda_lib&subdirectory=csrc/xentropy'`)"
-        )
-
-    try:
-        from flash_attn.ops.rms_norm import RMSNorm
-
-        class LlamaRMSNorm(RMSNorm):
-            """Patched LLamaRMSNorm"""
-
-            def __init__(self, hidden_size, eps=1e-6):
-                super().__init__(hidden_size, eps=eps)
-
-        LOG.info("patching with flash_attn.ops.rms_norm")
-        transformers.models.llama.modeling_llama.LlamaRMSNorm = LlamaRMSNorm
-    except ImportError:
-        LOG.info(
-            "optimized flash-attention RMSNorm not found (run `pip install 'git+https://github.com/Dao-AILab/flash-attention.git#egg=dropout_layer_norm&subdirectory=csrc/layer_norm'`)"
-        )
-

 # Disable the transformation of the attention mask in LlamaModel as the flash attention
 # requires the attention mask to be the same as the key_padding_mask
--- a/src/axolotl/prompters.py
+++ b/src/axolotl/prompters.py
@@ -309,6 +309,10 @@ class ShareGPTPrompter:  # pylint: disable=too-few-public-methods
        )

    def build_prompt(self, source) -> Generator[str, None, None]:
+        # ignore the system prompt if provided
+        if source[0]["from"] == "system":
+            source.pop(0)
+
        if len(source) < 2:
            # If there isn't a back and forth conversation, ignore it
            # also happens on the data splitting leaving empty conversations
@@ -317,12 +321,6 @@ class ShareGPTPrompter:  # pylint: disable=too-few-public-methods
            )

        conv = self._conversation.copy()
-
-        # Add the conversation system prompt if provided, otherwise use the default one
-        if source[0]["from"] == "system":
-            conv.system = source[0]["value"]
-            source.pop(0)
-
        roles = {"human": conv.roles[0], "gpt": conv.roles[1]}

        try:
--- a/src/axolotl/train.py
+++ b/src/axolotl/train.py
@@ -1,141 +0,0 @@
-"""Prepare and train a model on a dataset. Can also infer from a model or merge lora"""
-
-import logging
-import os
-import signal
-import sys
-from dataclasses import dataclass
-from pathlib import Path
-from typing import Optional
-
-import torch
-
-# add src to the pythonpath so we don't need to pip install this
-from datasets import Dataset
-from optimum.bettertransformer import BetterTransformer
-
-from axolotl.common.cli import TrainerCliArgs
-from axolotl.logging_config import configure_logging
-from axolotl.utils.dict import DictDefault
-from axolotl.utils.models import load_model, load_tokenizer
-from axolotl.utils.trainer import setup_trainer
-
-project_root = os.path.abspath(os.path.join(os.path.dirname(__file__), ".."))
-src_dir = os.path.join(project_root, "src")
-sys.path.insert(0, src_dir)
-
-configure_logging()
-LOG = logging.getLogger("axolotl.train")
-
-
-@dataclass
-class TrainDatasetMeta:
-    """
-    dataclass to capture the dataset specific options for training
-    """
-
-    train_dataset: Dataset
-    eval_dataset: Optional[Dataset] = None
-    total_num_steps: Optional[int] = None
-
-
-def train(
-    *,
-    cfg: DictDefault,
-    cli_args: TrainerCliArgs,
-    dataset_meta: TrainDatasetMeta,
-):
-    # load the tokenizer first
-    LOG.info(f"loading tokenizer... {cfg.tokenizer_config or cfg.base_model_config}")
-    tokenizer = load_tokenizer(cfg)
-
-    train_dataset = dataset_meta.train_dataset
-    eval_dataset = dataset_meta.eval_dataset
-    total_num_steps = dataset_meta.total_num_steps
-
-    # Load the model and tokenizer
-    LOG.info("loading model and (optionally) peft_config...")
-    model, peft_config = load_model(cfg, tokenizer, inference=cli_args.inference)
-
-    safe_serialization = cfg.save_safetensors is True
-
-    if cfg.resume_from_checkpoint is None and cfg.auto_resume_from_checkpoints:
-        possible_checkpoints = [
-            str(cp) for cp in Path(cfg.output_dir).glob("checkpoint-*")
-        ]
-        if len(possible_checkpoints) > 0:
-            sorted_paths = sorted(
-                possible_checkpoints,
-                key=lambda path: int(path.split("-")[-1]),
-            )
-            cfg.resume_from_checkpoint = sorted_paths[-1]
-            LOG.info(
-                f"Using Auto-resume functionality to start with checkpoint at {cfg.resume_from_checkpoint}"
-            )
-    resume_from_checkpoint = cfg.resume_from_checkpoint
-
-    trainer = setup_trainer(
-        cfg, train_dataset, eval_dataset, model, tokenizer, total_num_steps
-    )
-
-    model.config.use_cache = False
-
-    if torch.__version__ >= "2" and sys.platform != "win32":
-        LOG.info("Compiling torch model")
-        model = torch.compile(model)
-
-    # go ahead and presave, so we have the adapter config available to inspect
-    if peft_config:
-        LOG.info(f"Pre-saving adapter config to {cfg.output_dir}")
-        peft_config.save_pretrained(cfg.output_dir)
-    # additionally presave the tokenizer and model configs
-    if not Path(cfg.output_dir).is_dir():
-        os.makedirs(cfg.output_dir, exist_ok=True)
-    tokenizer.save_pretrained(str(Path(cfg.output_dir)))
-    model.config.save_pretrained(str(Path(cfg.output_dir)))
-
-    # In case we want to stop early with ctrl+c, this is a nice to have to save the pretrained model
-    if cfg.local_rank == 0:
-
-        def terminate_handler(_, __, model):
-            if cfg.flash_optimum:
-                model = BetterTransformer.reverse(model)
-            model.save_pretrained(cfg.output_dir, safe_serialization=safe_serialization)
-            sys.exit(0)
-
-        signal.signal(
-            signal.SIGINT, lambda signum, frame: terminate_handler(signum, frame, model)
-        )
-
-    LOG.info("Starting trainer...")
-    if cfg.group_by_length:
-        LOG.info("hang tight... sorting dataset for group_by_length")
-
-    if cfg.flash_optimum:
-        with torch.backends.cuda.sdp_kernel(
-            enable_flash=True, enable_math=True, enable_mem_efficient=True
-        ):
-            trainer.train(resume_from_checkpoint=resume_from_checkpoint)
-    else:
-        trainer.train(resume_from_checkpoint=resume_from_checkpoint)
-
-    LOG.info(f"Training Completed!!! Saving pre-trained model to {cfg.output_dir}")
-
-    if cfg.relora_steps:
-        if cfg.adapter == "lora" and not (cfg.load_in_4bit or cfg.load_in_8bit):
-            model = model.merge_and_unload()
-        else:
-            # final model weights have already been saved by `ReLoRACallback.on_train_end`
-            return model, tokenizer
-
-    # TODO do we need this fix? https://huggingface.co/docs/accelerate/usage_guides/fsdp#saving-and-loading
-    # only save on rank 0, otherwise it corrupts output on multi-GPU when multiple processes attempt to write the same file
-    if cfg.fsdp:
-        trainer.save_model(cfg.output_dir)
-    elif cfg.local_rank == 0:
-        if cfg.flash_optimum:
-            model = BetterTransformer.reverse(model)
-
-        model.save_pretrained(cfg.output_dir, safe_serialization=safe_serialization)
-
-    return model, tokenizer
--- a/src/axolotl/utils/callbacks.py
+++ b/src/axolotl/utils/callbacks.py
@@ -1,19 +1,9 @@
 """Callbacks for Trainer class"""

-from __future__ import annotations
-
 import logging
 import os
-from typing import TYPE_CHECKING, Dict, List

-import evaluate
-import numpy as np
-import pandas as pd
-import torch
-import torch.distributed as dist
-from datasets import load_dataset
 from optimum.bettertransformer import BetterTransformer
-from tqdm import tqdm
 from transformers import (
    TrainerCallback,
    TrainerControl,
@@ -23,21 +13,8 @@ from transformers import (
 from transformers.trainer_utils import PREFIX_CHECKPOINT_DIR, IntervalStrategy

 from axolotl.utils.bench import log_gpu_memory_usage
-from axolotl.utils.distributed import (
-    barrier,
-    broadcast_dict,
-    gather_scalar_from_all_ranks,
-    get_world_size,
-    is_distributed,
-    is_main_process,
-    zero_first,
-)
-
-if TYPE_CHECKING:
-    from axolotl.utils.trainer import AxolotlTrainingArguments

 LOG = logging.getLogger("axolotl.callbacks")
-IGNORE_INDEX = -100


 class SavePeftModelCallback(TrainerCallback):  # pylint: disable=too-few-public-methods
@@ -119,207 +96,3 @@ class GPUStatsCallback(
            log_gpu_memory_usage(LOG, "while training", self.cfg.device)
            self.logged = True
        return control
-
-
-def bench_eval_callback_factory(trainer, tokenizer):
-    accuracy = evaluate.load("accuracy")
-    abcd_idx = [
-        tokenizer("A", add_special_tokens=False).input_ids[0],
-        tokenizer("B", add_special_tokens=False).input_ids[0],
-        tokenizer("C", add_special_tokens=False).input_ids[0],
-        tokenizer("D", add_special_tokens=False).input_ids[0],
-        tokenizer("E", add_special_tokens=False).input_ids[0],
-        tokenizer("F", add_special_tokens=False).input_ids[0],
-        tokenizer("G", add_special_tokens=False).input_ids[0],
-    ]
-    bench_split = "eval"
-
-    def transform_bench_subject(example):
-        # Split on ':' and trim whitespace
-        parts = example["subject"].split(":")
-        first_part = (
-            parts[0].strip().lower().replace("-", "_")
-        )  # Lowercase the first part
-        second_part = (
-            parts[1].strip().replace("-", "_") if len(parts) > 1 else "all"
-        )  # Replace hyphens with underscores
-
-        # Return the transformed values
-        return {"name": first_part, "subject": second_part}
-
-    if trainer.args.bench_dataset == "mmlu-zs":
-        bench_dataset = load_dataset(
-            "openaccess-ai-collective/mmlu-evals",
-            data_files={
-                "eval": "zero_shot_mmlu_val.json",
-                "test": "zero_shot_mmlu_test.json",
-            },
-        )
-        # bench_dataset = bench_dataset.remove_columns("subject")
-    # MMLU Five-shot (Eval/Test only)
-    elif trainer.args.bench_dataset in ["mmlu", "mmlu-fs"]:
-        bench_dataset = load_dataset(
-            "openaccess-ai-collective/mmlu-evals",
-            data_files={
-                "eval": "five_shot_mmlu_val.json",
-                "test": "five_shot_mmlu_test.json",
-            },
-        )
-        # bench_dataset = bench_dataset.remove_columns('subject')
-    elif "/" in trainer.args.bench_dataset:
-        bench_ds = trainer.args.bench_dataset
-        bench_ds_name = "/".join(bench_ds.split("/", 2)[:2])
-        bench_ds_data_file = "/".join(bench_ds.split("/", 2)[2:])
-        bench_dataset = load_dataset(
-            bench_ds_name,
-            data_files={
-                "eval": bench_ds_data_file,
-            },
-        )
-        bench_dataset["eval"] = bench_dataset["eval"].map(transform_bench_subject)
-    else:
-        raise ValueError(
-            f"unhandled value `{trainer.args.bench_dataset}` for bench_dataset training args"
-        )
-    bench_dataset = bench_dataset[trainer.args.bench_split]
-    if trainer.args.max_bench_samples is not None:
-        bench_dataset = bench_dataset.select(range(trainer.args.max_bench_samples))
-
-    def tokenize_evals(example):
-        source = f"{tokenizer.bos_token}{example['input']}"
-        target = f"{example['output']}{tokenizer.eos_token}"
-
-        tokenized_source = tokenizer(
-            source,
-            max_length=2048,
-            truncation=True,
-            add_special_tokens=False,
-        )
-        tokenized_target = tokenizer(
-            target,
-            max_length=2048,
-            truncation=True,
-            add_special_tokens=False,
-        )
-        input_ids = tokenized_source["input_ids"] + tokenized_target["input_ids"]
-        labels = [IGNORE_INDEX] * len(tokenized_source["input_ids"]) + tokenized_target[
-            "input_ids"
-        ]
-
-        return {
-            "input_ids": input_ids,
-            "labels": labels,
-            "subject": example["subject"],
-        }
-
-    with zero_first(is_main_process()):
-        bench_dataset = bench_dataset.map(tokenize_evals)
-        bench_dataset = bench_dataset.filter(lambda x: x["labels"][-2] in abcd_idx)
-
-    class BenchEvalCallback(TrainerCallback):
-        """
-        TrainerCallback that runs the MMLU evals
-        """
-
-        def on_evaluate(
-            self,
-            args: AxolotlTrainingArguments,
-            state: TrainerState,  # pylint: disable=unused-argument
-            control: TrainerControl,  # pylint: disable=unused-argument
-            metrics: Dict[str, float],  # pylint: disable=unused-argument
-            **kwargs,  # pylint: disable=unused-argument
-        ):
-            data_loader = trainer.get_bench_dataloader(
-                bench_dataset.remove_columns(["input", "subject", "output", "name"])
-            )
-            trainer.model.eval()
-            preds, refs = [], []
-            loss_bench = 0
-            for batch in tqdm(data_loader, total=len(data_loader)):
-                (loss, logits, labels) = trainer.prediction_step(
-                    trainer.model,
-                    batch,
-                    prediction_loss_only=False,
-                )
-                # There are two tokens, the output, and eos token.
-                for i, logit in enumerate(logits):
-                    label_non_zero_id = (batch["labels"][i] != IGNORE_INDEX).nonzero()[
-                        0
-                    ][0]
-                    logit_abcd = logit[label_non_zero_id - 1][abcd_idx]
-                    preds.append(torch.argmax(logit_abcd).item())
-                labels = labels[labels != IGNORE_INDEX].view(-1, 2)[:, 0]
-                refs += [
-                    abcd_idx.index(label) if label in abcd_idx else -1
-                    for label in labels.tolist()
-                ]
-                loss_bench += loss.item()
-            # Extract results by subject.
-            bench_name = bench_dataset["name"]
-            bench_names: dict = {s: {"refs": [], "preds": []} for s in set(bench_name)}
-            for s, p, r in zip(bench_name, preds, refs):  # pylint: disable=invalid-name
-                bench_names[s]["preds"].append(p)
-                bench_names[s]["refs"].append(r)
-            barrier()
-            local_bench_names = bench_names
-            gathered_bench_names: List[Dict] = [{} for _ in range(get_world_size())]
-            # Gather results from all GPUs to GPU 0
-
-            loss_bench_ranks = gather_scalar_from_all_ranks(
-                lambda: loss_bench, get_world_size()
-            )
-            len_data_loader_ranks = gather_scalar_from_all_ranks(
-                lambda: len(data_loader), get_world_size()
-            )
-
-            results = {}
-            if is_distributed() and not is_main_process():
-                dist.gather_object(local_bench_names, dst=0)
-            else:
-                if is_distributed():
-                    dist.gather_object(local_bench_names, gathered_bench_names, dst=0)
-                else:
-                    gathered_bench_names = [local_bench_names]
-                bench_loss = sum(loss_bench_ranks) / sum(len_data_loader_ranks)
-                results = {f"{bench_split}_bench_loss": bench_loss}
-
-                # Combine results from all GPUs
-                combined_bench_names: Dict[str, Dict[str, List]] = {}
-                for bench_name in gathered_bench_names:
-                    for name, data in bench_name.items():
-                        if name not in combined_bench_names:
-                            combined_bench_names[name] = {"refs": [], "preds": []}
-                        combined_bench_names[name]["refs"].extend(data["refs"])
-                        combined_bench_names[name]["preds"].extend(data["preds"])
-
-                bench_scores = []
-                bench_refs = []
-                bench_preds = []
-                for (
-                    bench_name
-                ) in combined_bench_names:  # pylint: disable=consider-using-dict-items
-                    bench_score = accuracy.compute(
-                        references=combined_bench_names[bench_name]["refs"],
-                        predictions=combined_bench_names[bench_name]["preds"],
-                    )["accuracy"]
-                    bench_refs.extend(combined_bench_names[bench_name]["refs"])
-                    bench_preds.extend(combined_bench_names[bench_name]["preds"])
-                    if not pd.isna(bench_score):
-                        results[
-                            f"{bench_split}_bench_accuracy_{bench_name}"
-                        ] = bench_score
-                        bench_scores.append(bench_score)
-                    else:
-                        results[f"{bench_split}_bench_accuracy_{bench_name}"] = 0.0
-                        bench_scores.append(0.0)
-                results[f"{bench_split}_bench_average_accuracy"] = np.mean(bench_scores)
-                results[f"{bench_split}_bench_total_accuracy"] = accuracy.compute(
-                    references=bench_refs, predictions=bench_preds
-                )["accuracy"]
-                trainer.log(results)
-
-            results = broadcast_dict(results)
-            for key, val in results.items():
-                metrics[key] = val
-
-    return BenchEvalCallback
--- a/src/axolotl/utils/config.py
+++ b/src/axolotl/utils/config.py
@@ -6,7 +6,6 @@ import os
 import torch

 from axolotl.utils.bench import log_gpu_memory_usage
-from axolotl.utils.models import load_model_config

 LOG = logging.getLogger("axolotl")

@@ -70,16 +69,6 @@ def normalize_config(cfg):
    else:
        cfg.torch_dtype = torch.float32

-    model_config = load_model_config(cfg)
-
-    # figure out if the model is llama
-    cfg.is_llama_derived_model = (
-        (hasattr(model_config, "model_type") and model_config.model_type == "llama")
-        or cfg.is_llama_derived_model
-        or "llama" in cfg.base_model
-        or (cfg.model_type and "llama" in cfg.model_type.lower())
-    )
-
    log_gpu_memory_usage(LOG, "baseline", cfg.device)


@@ -97,11 +86,6 @@ def validate_config(cfg):
            )
        )

-    if cfg.sample_packing and not cfg.pad_to_sequence_len:
-        LOG.warning(
-            "`pad_to_sequence_len: true` is recommended when using sample_packing"
-        )
-
    if cfg.gradient_accumulation_steps and cfg.batch_size:
        raise ValueError(
            "please set only one of gradient_accumulation_steps or batch_size"
@@ -220,15 +204,6 @@ def validate_config(cfg):
            "sample_packing not compatible with xformers_attention. Use flash_attention"
        )

-    if cfg.early_stopping_patience:
-        if not cfg.save_steps or not cfg.eval_steps:
-            raise ValueError(
-                "`early_stopping_patience` requires save_steps and eval_steps to be set. eval_steps should evenly divide save_steps."
-            )
-        if cfg.save_steps % cfg.eval_steps != 0:
-            raise ValueError(
-                "`early_stopping_patience` requires that eval_steps should evenly divide save_steps."
-            )
    # TODO
    # MPT 7b
    # https://github.com/facebookresearch/bitsandbytes/issues/25
--- a/src/axolotl/utils/data.py
+++ b/src/axolotl/utils/data.py
@@ -2,6 +2,7 @@
 import functools
 import hashlib
 import logging
+from hashlib import md5
 from pathlib import Path
 from typing import Tuple, Union

@@ -51,13 +52,6 @@ LOG = logging.getLogger("axolotl")
 DEFAULT_DATASET_PREPARED_PATH = "last_run_prepared"


-def md5(to_hash: str, encoding: str = "utf-8") -> str:
-    try:
-        return hashlib.md5(to_hash.encode(encoding), usedforsecurity=False).hexdigest()
-    except TypeError:
-        return hashlib.md5(to_hash.encode(encoding)).hexdigest()  # nosec
-
-
 def prepare_dataset(cfg, tokenizer):
    if not cfg.pretraining_dataset:
        with zero_first(is_main_process()):
@@ -94,7 +88,7 @@ def load_tokenized_prepared_datasets(
 ) -> DatasetDict:
    tokenizer_name = tokenizer.__class__.__name__
    ds_hash = str(
-        md5(
+        md5(  # nosec
            (
                str(cfg.sequence_len)
                + "@"
@@ -103,8 +97,8 @@ def load_tokenized_prepared_datasets(
                )
                + "|"
                + tokenizer_name
-            )
-        )
+            ).encode("utf-8")
+        ).hexdigest()
    )
    prepared_ds_path = (
        Path(cfg.dataset_prepared_path) / ds_hash
@@ -380,7 +374,7 @@ def load_prepare_datasets(
        # see if we can go ahead and load the stacked dataset
        seed = f"@{str(cfg.seed)}" if cfg.seed else ""
        ds_hash = str(
-            md5(
+            md5(  # nosec
                (
                    str(cfg.sequence_len)
                    + "@"
@@ -391,8 +385,8 @@ def load_prepare_datasets(
                    )
                    + "|"
                    + tokenizer_name
-                )
-            )
+                ).encode("utf-8")
+            ).hexdigest()
        )
        prepared_ds_path = (
            Path(cfg.dataset_prepared_path) / ds_hash
@@ -506,8 +500,12 @@ def load_prepare_datasets(
            + "|"
            + str(cfg.seed or 42)
        )
-        train_fingerprint = md5(to_hash_train)
-        test_fingerprint = md5(to_hash_test)
+        train_fingerprint = hashlib.md5(
+            to_hash_train.encode(), usedforsecurity=False
+        ).hexdigest()
+        test_fingerprint = hashlib.md5(
+            to_hash_test.encode(), usedforsecurity=False
+        ).hexdigest()

        with zero_first(is_main_process()):
            dataset = dataset.train_test_split(
--- a/src/axolotl/utils/distributed.py
+++ b/src/axolotl/utils/distributed.py
@@ -1,11 +1,8 @@
 """
 utility helpers for distributed checks
 """
-import os
-import pickle  # nosec
 from contextlib import contextmanager

-import torch
 import torch.distributed as dist
 from accelerate import Accelerator

@@ -46,10 +43,6 @@ def is_main_process():
    return dist.get_rank() == 0


-def get_world_size():
-    return int(os.getenv("WORLD_SIZE", "1"))
-
-
@contextmanager
 def zero_first(is_main):
    """
@@ -60,64 +53,3 @@ def zero_first(is_main):
    yield
    if is_main:  # then rank 0 waits after it has run the context
        barrier()
-
-
-def gather_scalar_from_all_ranks(fn, world_size=1):  # pylint: disable=invalid-name
-    """
-    Run a callable 'fn' on all ranks and gather the results on the specified rank.
-
-    Args:
-    - fn (callable): A function that computes the value. This should not have any side effects.
-    - rank (int, optional): The rank that gathers the values. Default is 0.
-    - world_size (int, optional): Total number of processes in the current distributed setup.
-
-    Returns:
-    - A list of computed values from all ranks if on the gathering rank, otherwise None.
-    """
-    value_scalar = fn()
-    if not is_distributed():
-        return [value_scalar]
-    value_tensor = torch.tensor(value_scalar, device=dist.get_rank()).float()
-
-    if not is_main_process():
-        dist.gather(value_tensor, dst=0)
-    else:
-        gathered_tensors = [torch.zeros_like(value_tensor) for _ in range(world_size)]
-        dist.gather(value_tensor, gather_list=gathered_tensors, dst=0)
-
-        # Convert tensors back to their original type (int or float)
-        gathered_values = []
-        for tensor in gathered_tensors:
-            if tensor == tensor.int():
-                gathered_values.append(int(tensor.item()))
-            else:
-                gathered_values.append(float(tensor.item()))
-        return gathered_values
-    return None
-
-
-def broadcast_dict(vals: dict):
-    if not is_distributed():
-        return vals
-
-    if is_main_process():
-        data_byte = pickle.dumps(vals)
-        data_tensor = torch.ByteTensor(list(data_byte)).to("cuda")
-        data_size = torch.IntTensor([len(data_byte)]).to("cuda")
-    else:
-        data_tensor = torch.empty([1024], dtype=torch.uint8, device="cuda")
-        data_size = torch.IntTensor([0]).to("cuda")
-
-    dist.broadcast(data_size, 0)
-    if not is_main_process():
-        # resize
-        data_tensor = data_tensor.new_empty([data_size.item()])
-
-    dist.broadcast(data_tensor, 0)
-
-    if not is_main_process():
-        data_list = data_tensor.cpu().tolist()
-        data_byte = bytes(data_list[: data_size.item()])
-        vals = pickle.loads(data_byte)  # nosec
-
-    return vals
--- a/src/axolotl/utils/models.py
+++ b/src/axolotl/utils/models.py
@@ -159,13 +159,11 @@ def load_model(
    if cfg.model_revision:
        model_kwargs["revision"] = cfg.model_revision
    if cfg.gptq:
-        model_config = load_model_config(cfg)
-        if not hasattr(model_config, "quantization_config"):
-            LOG.warning("model config does not contain quantization_config information")
-        else:
-            model_kwargs["quantization_config"] = GPTQConfig(
-                **model_config.quantization_config
-            )
+        # TODO we should figure out how read the models config.json first
+        model_kwargs["quantization_config"] = GPTQConfig(
+            bits=cfg.gptq_bits,
+            disable_exllama=True,
+        )
    if cfg.adapter == "qlora" and cfg.load_in_4bit:
        model_kwargs["quantization_config"] = BitsAndBytesConfig(
            load_in_4bit=True,
@@ -328,7 +326,7 @@ def load_model(

    # LlamaRMSNorm layers are in fp32 after kbit_training or full finetune, so we need to
    # convert them back to fp16/bf16 for flash-attn compatibility.
-    if needs_fa2_dtype or (cfg.flash_attention and cfg.is_llama_derived_model):
+    if needs_fa2_dtype and (cfg.flash_attention and cfg.is_llama_derived_model):
        LOG.info("converting modules to %s for flash attention", cfg.torch_dtype)
        for name, module in model.named_modules():
            if "norm" in name:
--- a/src/axolotl/utils/tokenization.py
+++ b/src/axolotl/utils/tokenization.py
@@ -8,13 +8,13 @@ from termcolor import colored
 LOG = logging.getLogger("axolotl")


-def check_dataset_labels(dataset, tokenizer, num_examples=5, text_only=False):
+def check_dataset_labels(dataset, tokenizer):
    # the dataset is already shuffled, so let's just check the first 5 elements
-    for idx in range(num_examples):
-        check_example_labels(dataset[idx], tokenizer, text_only=text_only)
+    for idx in range(5):
+        check_example_labels(dataset[idx], tokenizer)


-def check_example_labels(example, tokenizer, text_only=False):
+def check_example_labels(example, tokenizer):
    # Get the input_ids, labels, and attention_mask from the dataset
    input_ids = example["input_ids"]
    labels = example["labels"]
@@ -29,10 +29,8 @@ def check_example_labels(example, tokenizer, text_only=False):
        decoded_input_token = tokenizer.decode(input_id)
        # Choose the color based on whether the label has the ignore value or not
        color = "red" if label_id == -100 else ("yellow" if label_id == 0 else "green")
-        colored_token = colored(decoded_input_token, color) + (
-            not text_only
-            and colored(f"({label_id}, {mask}, {input_id})", "white")
-            or ""
+        colored_token = colored(decoded_input_token, color) + colored(
+            f"({label_id}, {mask}, {input_id})", "white"
        )
        colored_tokens.append(colored_token)

--- a/src/axolotl/utils/trainer.py
+++ b/src/axolotl/utils/trainer.py
@@ -12,15 +12,9 @@ from typing import Optional, Union

 import numpy as np
 import torch.cuda
-import transformers
 from datasets import Dataset, set_caching_enabled
 from torch.optim.lr_scheduler import OneCycleLR
-from torch.utils.data import (
-    DataLoader,
-    DistributedSampler,
-    RandomSampler,
-    SequentialSampler,
-)
+from torch.utils.data import DataLoader, DistributedSampler, RandomSampler
 from transformers import EarlyStoppingCallback, Trainer, TrainingArguments
 from transformers.trainer_pt_utils import SequentialDistributedSampler

@@ -29,11 +23,9 @@ from axolotl.utils.callbacks import (
    GPUStatsCallback,
    SaveBetterTransformerModelCallback,
    SavePeftModelCallback,
-    bench_eval_callback_factory,
 )
 from axolotl.utils.collators import DataCollatorForSeq2Seq
 from axolotl.utils.dataloader import MultipackDistributedDataloader
-from axolotl.utils.distributed import is_main_process, zero_first
 from axolotl.utils.schedulers import get_cosine_schedule_with_quadratic_warmup

 LOG = logging.getLogger("axolotl")
@@ -135,27 +127,6 @@ class AxolotlTrainingArguments(TrainingArguments):
        default=None,
        metadata={"help": "how many warmup steps to take after reset for ReLoRA"},
    )
-    bench_split: Optional[str] = field(
-        default="eval", metadata={"help": "The benchmark split to run on"}
-    )
-    bench_dataset: Optional[str] = field(
-        default="pharaouk/dharma-1/dharma_1_mini.json",
-        metadata={
-            "help": "Benchmark dataset to use: options are `mmlu-zs`, `mmlu-fs`, or the full path to the dataset file"
-        },
-    )
-    do_bench_eval: Optional[bool] = field(
-        default=False, metadata={"help": "Whether to run the Benchmark evaluation."}
-    )
-    max_bench_samples: Optional[int] = field(
-        default=None,
-        metadata={
-            "help": "If set, only evaluates on `max_bench_samples` of the benchmark dataset."
-        },
-    )
-    bench_source_max_len: int = field(
-        default=2048, metadata={"help": "Maximum source sequence length for bench."}
-    )


 class AxolotlTrainer(Trainer):
@@ -165,10 +136,6 @@ class AxolotlTrainer(Trainer):

    args = None  # type: AxolotlTrainingArguments

-    def __init__(self, *args, bench_data_collator=None, **kwargs):
-        self.bench_data_collator = bench_data_collator
-        super().__init__(*args, **kwargs)
-
    def create_scheduler(
        self, num_training_steps: int, optimizer: torch.optim.Optimizer = None
    ):
@@ -259,31 +226,6 @@ class AxolotlTrainer(Trainer):
            )
        return super().get_eval_dataloader(eval_dataset)

-    def _get_bench_sampler(
-        self, bench_dataset: Dataset
-    ) -> Optional[torch.utils.data.Sampler]:
-        if self.args.world_size <= 1:
-            return SequentialSampler(bench_dataset)
-        return None
-
-    def get_bench_dataloader(
-        self,
-        bench_dataset: Dataset,
-    ) -> Union[DataLoader, MultipackDistributedDataloader]:
-        dataloader_params = {
-            "batch_size": self.args.eval_batch_size,
-            "collate_fn": self.bench_data_collator,
-            "num_workers": self.args.dataloader_num_workers,
-            "pin_memory": self.args.dataloader_pin_memory,
-        }
-
-        if not isinstance(bench_dataset, torch.utils.data.IterableDataset):
-            dataloader_params["sampler"] = self._get_bench_sampler(bench_dataset)
-            dataloader_params["drop_last"] = self.args.dataloader_drop_last
-
-        return DataLoader(bench_dataset, **dataloader_params)
-        # return self.accelerator.prepare(DataLoader(bench_dataset, **dataloader_params))
-
    def compute_loss(self, model, inputs, return_outputs=False):
        # use one's weighted cross entropy loss calc
        # if self.args.sample_packing:
@@ -362,7 +304,7 @@ def add_position_ids(sample):


 def drop_long_seq(sample, sequence_len=2048):
-    return len(sample["input_ids"]) <= sequence_len and len(sample["input_ids"]) > 0
+    return len(sample["input_ids"]) <= sequence_len


@contextmanager
@@ -376,17 +318,14 @@ def disable_datasets_caching():

 def process_datasets_for_packing(cfg, train_dataset, eval_dataset):
    drop_long = partial(drop_long_seq, sequence_len=cfg.sequence_len)
-    with zero_first(is_main_process()):
-        train_dataset = train_dataset.filter(drop_long, num_proc=os.cpu_count())
-        if eval_dataset:
-            eval_dataset = eval_dataset.filter(drop_long, num_proc=os.cpu_count())
+    train_dataset = train_dataset.filter(drop_long, num_proc=os.cpu_count())
+    if eval_dataset:
+        eval_dataset = eval_dataset.filter(drop_long, num_proc=os.cpu_count())

-        if cfg.sample_packing:
-            train_dataset = train_dataset.map(add_position_ids, num_proc=os.cpu_count())
-            if eval_dataset:
-                eval_dataset = eval_dataset.map(
-                    add_position_ids, num_proc=os.cpu_count()
-                )
+    if cfg.sample_packing:
+        train_dataset = train_dataset.map(add_position_ids, num_proc=os.cpu_count())
+        if eval_dataset:
+            eval_dataset = eval_dataset.map(add_position_ids, num_proc=os.cpu_count())
    return train_dataset, eval_dataset


@@ -405,16 +344,6 @@ def calculate_total_num_steps(cfg, train_dataset, tokenizer):
            LOG.info(f"📝 UPDATE CONFIG WITH: `total_num_tokens: {total_num_tokens}`")
            cfg.total_num_tokens = total_num_tokens

-        if not cfg.total_supervised_tokens:
-            total_supervised_tokens = (
-                train_dataset.data.column("labels")
-                .to_pandas()
-                .apply(lambda x: np.sum(np.array(x) != -100))
-                .sum()
-            )
-            LOG.info(f"`total_supervised_tokens: {total_supervised_tokens}`")
-            cfg.total_supervised_tokens = total_supervised_tokens
-
        if cfg.sample_packing_eff_est:
            total_num_steps = (
                # match count to len est in dataloader
@@ -572,24 +501,6 @@ def setup_trainer(cfg, train_dataset, eval_dataset, model, tokenizer, total_num_
            "steps" if cfg.save_steps else "epoch"
        )

-    if cfg.do_bench_eval:
-        training_arguments_kwargs["do_bench_eval"] = cfg.do_bench_eval
-        if cfg.bench_dataset:
-            training_arguments_kwargs["bench_dataset"] = cfg.bench_dataset
-    if cfg.metric_for_best_model:
-        training_arguments_kwargs["metric_for_best_model"] = cfg.metric_for_best_model
-    if cfg.greater_is_better:
-        training_arguments_kwargs["greater_is_better"] = cfg.greater_is_better
-
-    # DDP Config
-    if cfg.ddp_timeout:
-        training_arguments_kwargs["ddp_timeout"] = cfg.ddp_timeout
-    # see https://pytorch.org/docs/stable/generated/torch.nn.parallel.DistributedDataParallel.html
-    if cfg.ddp_bucket_cap_mb:
-        training_arguments_kwargs["ddp_bucket_cap_mb"] = cfg.ddp_bucket_cap_mb
-    if cfg.ddp_broadcast_buffers is not None:
-        training_arguments_kwargs["ddp_broadcast_buffers"] = cfg.ddp_broadcast_buffers
-
    training_args = AxolotlTrainingArguments(  # pylint: disable=unexpected-keyword-arg
        max_steps=total_num_steps if cfg.max_steps else -1,
        max_seq_length=cfg.sequence_len,
@@ -605,10 +516,11 @@ def setup_trainer(cfg, train_dataset, eval_dataset, model, tokenizer, total_num_
        output_dir=cfg.output_dir,
        save_total_limit=cfg.save_total_limit if cfg.save_total_limit else 4,
        load_best_model_at_end=(
-            (cfg.load_best_model_at_end is not False or cfg.early_stopping_patience)
+            cfg.load_best_model_at_end is not False
            and cfg.val_set_size > 0
            and cfg.save_steps
            and cfg.save_steps % cfg.eval_steps == 0
+            and cfg.load_in_8bit is not True
        )
        or False,
        ddp_find_unused_parameters=False if cfg.ddp else None,
@@ -640,6 +552,13 @@ def setup_trainer(cfg, train_dataset, eval_dataset, model, tokenizer, total_num_
    if cfg.relora_steps:
        callbacks.append(ReLoRACallback(cfg))

+    # TODO on_save callback to sync checkpoints to GCP/AWS in background
+    if cfg.early_stopping_patience:
+        early_stop_cb = EarlyStoppingCallback(
+            cfg.early_stopping_patience,
+        )
+        callbacks.append(early_stop_cb)
+
    if cfg.local_rank == 0 and cfg.adapter in [
        "lora",
        "qlora",
@@ -694,23 +613,8 @@ def setup_trainer(cfg, train_dataset, eval_dataset, model, tokenizer, total_num_
            return_tensors="pt",
            **data_collator_kwargs,
        ),
-        bench_data_collator=transformers.DataCollatorForSeq2Seq(
-            tokenizer,
-            return_tensors="pt",
-            **data_collator_kwargs,
-        ),
        callbacks=callbacks,
        **trainer_kwargs,
    )

-    if cfg.do_bench_eval:
-        trainer.add_callback(bench_eval_callback_factory(trainer, tokenizer))
-
-    # TODO on_save callback to sync checkpoints to GCP/AWS in background
-    if cfg.early_stopping_patience:
-        early_stop_cb = EarlyStoppingCallback(
-            cfg.early_stopping_patience,
-        )
-        trainer.add_callback(early_stop_cb)
-
    return trainer
--- a/tests/test_data.py
+++ b/tests/test_data.py
@@ -1,64 +0,0 @@
-"""
-test module for the axolotl.utis.data module
-"""
-import unittest
-
-from transformers import LlamaTokenizer
-
-from axolotl.utils.data import encode_pretraining, md5
-
-
-class TestEncodePretraining(unittest.TestCase):
-    """
-    test class for encode pretraining and md5 helper
-    """
-
-    def setUp(self):
-        self.tokenizer = LlamaTokenizer.from_pretrained("huggyllama/llama-7b")
-        self.tokenizer.add_special_tokens(
-            {
-                "eos_token": "</s>",
-                "bos_token": "<s>",
-                "unk_token": "<unk>",
-                "pad_token": "<pad>",
-            }
-        )
-        self.max_tokens = 15  # set a small number for easy inspection
-
-    def test_encode_pretraining(self):
-        examples = {
-            "text": [
-                "Hello, world!",
-                "Nice to meet you.",
-                "lorem ipsum dolor sit amet.",
-                "Nice to meet you again!.",
-                "hello, hello",
-            ]
-        }
-        result = encode_pretraining(self.tokenizer, self.max_tokens, examples)
-
-        self.assertEqual(len(result["input_ids"]), 3)
-
-        # Assert the length of input_ids and attention_mask is correct
-        self.assertEqual(len(result["input_ids"][0]), self.max_tokens)
-        self.assertEqual(len(result["attention_mask"][0]), self.max_tokens)
-
-        # Assert EOS and PAD tokens are correctly added
-        # hello world! is 4 tokens
-        self.assertEqual(result["input_ids"][0][0], self.tokenizer.bos_token_id)
-        self.assertEqual(result["input_ids"][0][5], self.tokenizer.eos_token_id)
-        self.assertEqual(result["input_ids"][0][6], self.tokenizer.pad_token_id)
-        # second part, 5 tokens
-        self.assertEqual(result["input_ids"][0][7], self.tokenizer.bos_token_id)
-        self.assertEqual(result["input_ids"][0][13], self.tokenizer.eos_token_id)
-        self.assertEqual(result["input_ids"][0][14], self.tokenizer.pad_token_id)
-
-    def test_md5(self):
-        self.assertEqual(md5("hello world"), "5eb63bbbe01eeed093cb22bb8f5acdc3")
-        self.assertEqual(
-            md5("hello world", "utf-8"), "5eb63bbbe01eeed093cb22bb8f5acdc3"
-        )
-
-
-if __name__ == "__main__":
-    unittest.main()
--- a/tests/test_validation.py
+++ b/tests/test_validation.py
@@ -328,20 +328,6 @@ class ValidationTest(unittest.TestCase):
                for record in self._caplog.records
            )

-        cfg = DictDefault(
-            {
-                "sample_packing": True,
-                "pad_to_sequence_len": None,
-            }
-        )
-        with self._caplog.at_level(logging.WARNING):
-            validate_config(cfg)
-            assert any(
-                "`pad_to_sequence_len: true` is recommended when using sample_packing"
-                in record.message
-                for record in self._caplog.records
-            )
-
        cfg = DictDefault(
            {
                "max_packed_sequence_len": 2048,
Author	SHA1	Message	Date
Wing Lian	0026fcc3df	remove torch install for now Some checks failed pre-commit / pre-commit (push) Has been cancelled Details PyTest / test (3.10) (push) Has been cancelled Details PyTest / test (3.9) (push) Has been cancelled Details	2023-09-01 08:15:22 -07:00
Wing Lian	b448c77148	address pr feedback	2023-08-29 22:45:22 -07:00
Wing Lian	c820d04669	gptq doesn't play well with sample packing	2023-08-29 12:15:31 -07:00
Wing Lian	588cd65a64	fix setup.py to use extra index url install torch for tests fix cuda version for autogptq index set torch in requirements so that it installs properly move gptq install around to work with github cicd	2023-08-29 12:02:51 -07:00
Wing Lian	caa80e891d	don't need explicit peft install for tests	2023-08-29 12:02:51 -07:00
Wing Lian	ac37753aa2	remove old gptq docker	2023-08-29 12:02:50 -07:00
Wing Lian	a29560004b	more tweaks and add yml	2023-08-29 12:02:00 -07:00
Wing Lian	1deb767fe8	auto gptq support	2023-08-29 11:31:14 -07:00