fix the context manager call

start to swap out for accelerate partial state
2023-09-03 22:49:28 -04:00 · 2023-09-03 22:41:29 -04:00
34 changed files with 306 additions and 552 deletions
--- a/.github/workflows/main.yml
+++ b/.github/workflows/main.yml
@@ -23,6 +23,11 @@ jobs:
            python_version: "3.10"
            pytorch: 2.0.1
            axolotl_extras:
+          - cuda: 118
+            cuda_version: 11.8.0
+            python_version: "3.9"
+            pytorch: 2.0.1
+            axolotl_extras: gptq
    runs-on: self-hosted
    steps:
      - name: Checkout
@@ -68,6 +73,11 @@ jobs:
            pytorch: 2.0.1
            axolotl_extras:
            is_latest: true
+          - cuda: 118
+            cuda_version: 11.8.0
+            python_version: "3.9"
+            pytorch: 2.0.1
+            axolotl_extras: gptq
    runs-on: self-hosted
    steps:
      - name: Checkout
--- a/.github/workflows/pypi.yml
+++ b/.github/workflows/pypi.yml
@@ -1,45 +0,0 @@
-name: publish pypi
-
-on:
-  push:
-    tags:
-      - '*'
-
-jobs:
-  pypi-publish:
-    name: Upload release to PyPI
-    runs-on: ubuntu-latest
-    environment:
-      name: pypi
-      url: https://pypi.org/p/axolotl
-    permissions:
-      id-token: write  # IMPORTANT: this permission is mandatory for trusted publishing
-    steps:
-      - name: Check out repository code
-        uses: actions/checkout@v3
-
-      - name: Setup Python
-        uses: actions/setup-python@v4
-        with:
-          python-version: "3.10"
-
-      - name: Install dependencies
-        run: |
-          pip3 install wheel
-          pip3 install -e .
-          pip3 install -r requirements-tests.txt
-
-      - name: Extract tag name
-        id: tag
-        run: echo ::set-output name=TAG_NAME::$(echo $GITHUB_REF | cut -d / -f 3)
-
-      - name: Update version in setup.py
-        run: >-
-          sed -i -E 's/version="([0-9.]+)",/version="${{ steps.tag.outputs.TAG_NAME }}",/g' setup.py
-
-      - name: Build a binary wheel
-        run: >-
-          python setup.py sdist bdist_wheel
-
-      - name: Publish package distributions to PyPI
-        uses: pypa/gh-action-pypi-publish@release/v1
--- a/.github/workflows/tests.yml
+++ b/.github/workflows/tests.yml
@@ -24,8 +24,8 @@ jobs:

      - name: Install dependencies
        run: |
-          pip3 install -e .
-          pip3 install -r requirements-tests.txt
+          pip install -e .[peft]
+          pip install -r requirements-tests.txt

      - name: Run tests
        run: |
--- a/README.md
+++ b/README.md
@@ -90,7 +90,8 @@ accelerate launch scripts/finetune.py examples/openllama-3b/lora.yml \
  ```bash
  docker run --gpus '"all"' --rm -it winglian/axolotl:main-py3.10-cu118-2.0.1
  ```
-  - `winglian/axolotl-runpod:main-latest`: for runpod or use this [direct link](https://runpod.io/gsc?template=v2ickqhz9s&ref=6i7fkpdz)
+  - `winglian/axolotl-runpod:main-py3.10-cu118-2.0.1`: for runpod
+  - `winglian/axolotl-runpod:main-py3.9-cu118-2.0.1-gptq`: for gptq

  Or run on the current files for development:

@@ -103,9 +104,19 @@ accelerate launch scripts/finetune.py examples/openllama-3b/lora.yml \

  2. Install pytorch stable https://pytorch.org/get-started/locally/

-  3. Install axolotl along with python dependencies
+  3. Install python dependencies with ONE of the following:
+      - Recommended, supports QLoRA, NO gptq/int4 support
        ```bash
-        pip3 install -e .[flash-attn]
+        pip3 install -e .
+        pip3 install -U git+https://github.com/huggingface/peft.git
+        ```
+      - gptq/int4 support, NO QLoRA
+        ```bash
+        pip3 install -e .[gptq]
+        ```
+      - same as above but not recommended
+        ```bash
+        pip3 install -e .[gptq_triton]
        ```

 - LambdaLabs
@@ -140,9 +151,10 @@ accelerate launch scripts/finetune.py examples/openllama-3b/lora.yml \
  git clone https://github.com/OpenAccess-AI-Collective/axolotl
  cd axolotl

-  pip3 install -e .
+  pip3 install -e . # change depend on needs
  pip3 install protobuf==3.20.3
  pip3 install -U --ignore-installed requests Pillow psutil scipy
+  pip3 install git+https://github.com/huggingface/peft.git # not for gptq
  ```

  5. Set path
@@ -560,30 +572,6 @@ log_sweep_min_lr:
 log_sweep_max_lr:

 # specify optimizer
-# Valid values are driven by the Transformers OptimizerNames class, see:
-# https://github.com/huggingface/transformers/blob/95b374952dc27d8511541d6f5a4e22c9ec11fb24/src/transformers/training_args.py#L134
-#
-# Note that not all optimizers may be available in your environment, ex: 'adamw_anyprecision' is part of
-# torchdistx, 'adamw_bnb_8bit' is part of bnb.optim.Adam8bit, etc. When in doubt, it is recommended to start with the optimizer used
-# in the examples/ for your model and fine-tuning use case.
-#
-# Valid values for 'optimizer' include:
-# - adamw_hf
-# - adamw_torch
-# - adamw_torch_fused
-# - adamw_torch_xla
-# - adamw_apex_fused
-# - adafactor
-# - adamw_anyprecision
-# - sgd
-# - adagrad
-# - adamw_bnb_8bit
-# - lion_8bit
-# - lion_32bit
-# - paged_adamw_32bit
-# - paged_adamw_8bit
-# - paged_lion_32bit
-# - paged_lion_8bit
 optimizer:
 # specify weight decay
 weight_decay:
@@ -764,10 +752,6 @@ Try to turn off xformers.

 It's safe to ignore it.

-> NCCL Timeouts during training
-
-See the [NCCL](docs/nccl.md) guide.
-
 ## Need help? 🙋♂️

 Join our [Discord server](https://discord.gg/HhrNrHJPRb) where we can help you
--- a/docker-compose.yaml
+++ b/docker-compose.yaml
@@ -9,11 +9,6 @@ services:
      - ~/.cache/huggingface/:/root/.cache/huggingface/
    # set environment variables
    environment:
-      # Set environment variables
-      - GIT_AUTHOR_NAME=${GIT_AUTHOR_NAME}
-      - GIT_AUTHOR_EMAIL=${GIT_AUTHOR_EMAIL}
-      - GIT_COMMITTER_NAME=${GIT_COMMITTER_NAME}
-      - GIT_COMMITTER_EMAIL=${GIT_COMMITTER_EMAIL}
      - WANDB_API_KEY=${WANDB_API_KEY}
    deploy:
      resources:
--- a/docker/Dockerfile
+++ b/docker/Dockerfile
@@ -11,6 +11,7 @@ RUN apt-get update && \

 WORKDIR /workspace

+RUN pip3 install "peft @ git+https://github.com/huggingface/peft.git@main"
 RUN git clone --depth=1 https://github.com/OpenAccess-AI-Collective/axolotl.git
 # If AXOLOTL_EXTRAS is set, append it in brackets
 RUN cd axolotl && \
--- a/docs/nccl.md
+++ b/docs/nccl.md
@@ -1,46 +0,0 @@
-# NCCL
-
-NVIDIA NCCL is a library to facilitate and optimize multi-GPU communication operations, such as broadcast, all-gather, reduce, all-reduce, etc. Broadly, NCCL configuration is highly environment-specific and is configured via several [environment variables](https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/env.html). A common NCCL-related problem occurs when a long-running operation times out causing the training process to abort:
-
-```text
-Watchdog caught collective operation timeout: WorkNCCL(SeqNum=42, OpType=ALLGATHER, Timeout(ms)=1800000) ran for 1806948 milliseconds before timing out.
-```
-
-Often, this timeout will happen after 30 minutes (the default setting) and is accompanied by below-average power consumption with near 100% GPU utilization before the error is raised. Nvidia recommends [disabling PCI access control services (ACS)](https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/troubleshooting.html#pci-access-control-services-acs) as a possible solution if this is available to you.
-
-Forcing cross-GPU communication via [NVLink](https://en.wikipedia.org/wiki/NVLink) may help without increasing timeouts. To verify that your configuration is leveraging NVLink run the following command:
-
-```shell
-nvidia-smi nvlink --status
-```
-
-To force NCCL to use NVLink, simply set this in the environment:
-
-```shell
-export NCCL_P2P_LEVEL=NVL
-```
-
-If NVLink is not available in your environment there are other options for ``NCCL_P2P_LEVEL`` in the table below:
-
-| NCCL_P2P_LEVEL | Description |
-| -------------- | ----------- |
-| PIX | P2P data transfers through no more than a single PCIe bridge. Faster data transfer rates vs to paths involving multiple bridges, but slower compared to direct GPU-to-GPU communication. |
-| PXB | P2P data transfers through multiple PCIe bridges but not going through the PCIe Host Bridge; this path involves a complex routing process, potentially incurring a moderate level of latency. |
-| PHB | P2P data transfers occur over the PCIe and through a PCIe Host Bridge, typically involving the CPU, which can facilitate direct memory access but might introduce additional latency compared to more direct paths (ex PIX, NVL) |
-
-To validate that acceptable data transfer speeds exist for your training job, running [NCCL Tests](https://github.com/NVIDIA/nccl-tests/blob/master/README.md) can help pinpoint bottlenecks, for example:
-
-```shell
-./build/all_reduce_perf -b 8 -e 128M -f 2 -g 3
-```
-
-It can be useful when debugging NCCL communication timeouts to activate additional logging in both PyTorch and NCCL:
-
-```shell
-export NCCL_DEBUG=INFO
-export NCCL_DEBUG_SUBSYS=ALL
-export TORCH_DISTRIBUTED_DEBUG=INFO
-export TORCHELASTIC_ERROR_FILE=/PATH/TO/torcherror.log
-```
-
-Finally, if you believe your training job needs more time you can increase the timeout past 30 minutes by setting the ``ddp_timeout`` value in the Axolotl configuration. See [PyTorch init_process_group](https://pytorch.org/docs/stable/distributed.html#torch.distributed.init_process_group) for documentation on this value.
--- a/examples/code-llama/13b/lora.yml
+++ b/examples/code-llama/13b/lora.yml
@@ -17,7 +17,6 @@ output_dir: ./lora-out

 sequence_len: 100000
 sample_packing: true
-pad_to_sequence_len: true

 adapter: lora
 lora_model_dir:
--- a/examples/code-llama/13b/qlora.yml
+++ b/examples/code-llama/13b/qlora.yml
@@ -20,7 +20,6 @@ lora_model_dir:

 sequence_len: 100000
 sample_packing: true
-pad_to_sequence_len: true

 lora_r: 32
 lora_alpha: 16
--- a/examples/code-llama/34b/lora.yml
+++ b/examples/code-llama/34b/lora.yml
@@ -17,7 +17,6 @@ output_dir: ./lora-out

 sequence_len: 100000
 sample_packing: true
-pad_to_sequence_len: true

 adapter: lora
 lora_model_dir:
--- a/examples/code-llama/34b/qlora.yml
+++ b/examples/code-llama/34b/qlora.yml
@@ -20,7 +20,6 @@ lora_model_dir:

 sequence_len: 100000
 sample_packing: true
-pad_to_sequence_len: true

 lora_r: 32
 lora_alpha: 16
--- a/examples/code-llama/7b/lora.yml
+++ b/examples/code-llama/7b/lora.yml
@@ -17,7 +17,6 @@ output_dir: ./lora-out

 sequence_len: 100000
 sample_packing: true
-pad_to_sequence_len: true

 adapter: lora
 lora_model_dir:
--- a/examples/code-llama/7b/qlora.yml
+++ b/examples/code-llama/7b/qlora.yml
@@ -20,7 +20,6 @@ lora_model_dir:

 sequence_len: 100000
 sample_packing: true
-pad_to_sequence_len: true

 lora_r: 32
 lora_alpha: 16
--- a/examples/gptq-lora-7b/README.md
+++ b/examples/gptq-lora-7b/README.md
@@ -0,0 +1,8 @@
+# LLaMa 7B using LoRA
+
+This is a good place to start for beginners. This will run on an NVIDIA RTX4090 with no other changes needed.
+
+```shell
+accelerate launch scripts/finetune.py examples/gptq-lora-7b/config.yml
+
+```
--- a/examples/gptq-lora-7b/config.yml
+++ b/examples/gptq-lora-7b/config.yml
@@ -0,0 +1,63 @@
+base_model: Neko-Institute-of-Science/LLaMA-7B-4bit-128g
+base_model_config: Neko-Institute-of-Science/LLaMA-7B-4bit-128g
+model_type: LlamaForCausalLM
+tokenizer_type: LlamaTokenizer
+trust_remote_code:
+load_in_8bit: true
+gptq: true
+datasets:
+  - path: vicgalle/alpaca-gpt4
+    type: alpaca
+dataset_prepared_path: last_run_prepared
+val_set_size: 0.02
+adapter:
+lora_model_dir:
+sequence_len: 2048
+max_packed_sequence_len:
+lora_r: 8
+lora_alpha: 16
+lora_dropout: 0.05
+lora_target_modules:
+  - q_proj
+  - v_proj
+lora_fan_in_fan_out: false
+wandb_project: llama-7b-lora-int4
+wandb_entity:
+wandb_watch:
+wandb_run_id:
+wandb_log_model:
+output_dir: ./llama-7b-lora-int4
+gradient_accumulation_steps: 1
+micro_batch_size: 1
+num_epochs: 3
+optimizer: adamw_bnb_8bit
+torchdistx_path:
+lr_scheduler: cosine
+learning_rate: 0.0000002
+train_on_inputs: false
+group_by_length: false
+fp16: true
+bf16: false
+tf32: true
+early_stopping_patience:
+resume_from_checkpoint:
+local_rank:
+logging_steps: 5
+xformers_attention:
+flash_attention:
+gradient_checkpointing: true
+gptq_groupsize: 128
+gptq_model_v1: false
+warmup_steps: 20
+eval_steps: 110
+save_steps: 660
+debug:
+deepspeed:
+weight_decay: 0.0001
+fsdp:
+fsdp_config:
+tokens:
+  pad_token: "<pad>"
+  bos_token: "<s>"
+  eos_token: "</s>"
+  unk_token: "<unk>"
--- a/examples/llama-2/gptq-lora.yml
+++ b/examples/llama-2/gptq-lora.yml
@@ -1,76 +0,0 @@
-base_model: TheBloke/Llama-2-7B-GPTQ
-base_model_config: TheBloke/Llama-2-7B-GPTQ
-is_llama_derived_model: false
-gptq: true
-gptq_bits: 4
-model_type: AutoModelForCausalLM
-tokenizer_type: LlamaTokenizer
-tokenizer_use_fast: true
-tokenizer_legacy: true
-load_in_8bit: false
-load_in_4bit: false
-strict: false
-push_dataset_to_hub:
-hf_use_auth_token: true
-datasets:
-  - path: mhenrichsen/alpaca_2k_test
-    type: alpaca
-dataset_prepared_path: last_run_prepared
-val_set_size: 0.01
-adapter: lora
-lora_model_dir:
-sequence_len: 4096
-sample_packing:
-lora_r: 8
-lora_alpha: 32
-lora_dropout: 0.05
-lora_target_modules:
-  - k_proj
-  - o_proj
-  - q_proj
-  - v_proj
-lora_target_linear:
-lora_fan_in_fan_out:
-wandb_project:
-wandb_watch:
-wandb_run_id:
-wandb_log_model:
-output_dir: ./model-out
-gradient_accumulation_steps: 1
-micro_batch_size: 1
-num_epochs: 3
-optimizer: adamw_torch
-adam_beta2: 0.95
-adam_eps: 0.00001
-max_grad_norm: 1.0
-torchdistx_path:
-lr_scheduler: cosine
-lr_quadratic_warmup: true
-learning_rate: 0.000017
-train_on_inputs: false
-group_by_length: false
-bf16: false
-fp16: false
-float16: true
-tf32: true
-gradient_checkpointing: true
-early_stopping_patience:
-resume_from_checkpoint:
-local_rank:
-logging_steps: 1
-xformers_attention:
-flash_attention:
-sdp_attention:
-flash_optimum:
-gptq_groupsize:
-gptq_model_v1:
-warmup_steps: 100
-eval_steps:
-save_steps:
-debug:
-deepspeed:
-weight_decay: 0.1
-special_tokens:
-  bos_token: "<s>"
-  eos_token: "</s>"
-  unk_token: "<unk>"
--- a/examples/llama-2/lora.yml
+++ b/examples/llama-2/lora.yml
@@ -17,7 +17,6 @@ output_dir: ./lora-out

 sequence_len: 4096
 sample_packing: true
-pad_to_sequence_len: true

 adapter: lora
 lora_model_dir:
--- a/examples/llama-2/qlora.yml
+++ b/examples/llama-2/qlora.yml
@@ -20,7 +20,6 @@ lora_model_dir:

 sequence_len: 4096
 sample_packing: true
-pad_to_sequence_len: true

 lora_r: 32
 lora_alpha: 16
--- a/examples/llama-2/relora.yml
+++ b/examples/llama-2/relora.yml
@@ -20,7 +20,6 @@ lora_model_dir:

 sequence_len: 4096
 sample_packing: true
-pad_to_sequence_len: true

 lora_r: 8
 lora_alpha: 16
--- a/requirements.txt
+++ b/requirements.txt
@@ -1,18 +1,14 @@
--extra-index-url https://download.pytorch.org/whl/cu118
--extra-index-url https://huggingface.github.io/autogptq-index/whl/cu118/
-torch==2.0.1
-auto-gptq
 packaging
 peft @ git+https://github.com/huggingface/peft.git
 transformers @ git+https://github.com/huggingface/transformers.git
 bitsandbytes>=0.41.1
-accelerate @ git+https://github.com/huggingface/accelerate
+accelerate @ git+https://github.com/huggingface/accelerate@2a289f6108e77a77a4efffb3f6316bc98538413b
 addict
 evaluate
 fire
 PyYAML>=6.0
 datasets
-flash-attn>=2.2.1
+flash-attn>=2.0.8
 sentencepiece
 wandb
 einops
--- a/scripts/finetune.py
+++ b/scripts/finetune.py
@@ -24,7 +24,7 @@ from axolotl.utils.config import normalize_config, validate_config
 from axolotl.utils.data import prepare_dataset
 from axolotl.utils.dict import DictDefault
 from axolotl.utils.distributed import is_main_process
-from axolotl.utils.models import load_tokenizer
+from axolotl.utils.models import load_model_config, load_tokenizer
 from axolotl.utils.tokenization import check_dataset_labels
 from axolotl.utils.wandb import setup_wandb_env_vars

@@ -216,6 +216,15 @@ def load_cfg(config: Path = Path("examples/"), **kwargs):
            else:
                cfg[k] = kwargs[k]

+    model_config = load_model_config(cfg)
+
+    # figure out if the model is llama
+    cfg.is_llama_derived_model = (
+        (hasattr(model_config, "model_type") and model_config.model_type == "llama")
+        or cfg.is_llama_derived_model
+        or "llama" in cfg.base_model
+        or (cfg.model_type and "llama" in cfg.model_type.lower())
+    )
    validate_config(cfg)

    normalize_config(cfg)
--- a/setup.py
+++ b/setup.py
@@ -2,41 +2,38 @@

 from setuptools import find_packages, setup

-
-def parse_requirements():
-    _install_requires = []
-    _dependency_links = []
-    with open("./requirements.txt", encoding="utf-8") as requirements_file:
-        lines = [r.strip() for r in requirements_file.readlines()]
-        for line in lines:
-            if line.startswith("--extra-index-url"):
-                # Handle custom index URLs
-                _, url = line.split()
-                _dependency_links.append(url)
-            elif "flash-attn" not in line and line and line[0] != "#":
-                # Handle standard packages
-                _install_requires.append(line)
-    return _install_requires, _dependency_links
-
-
-install_requires, dependency_links = parse_requirements()
-
+install_requires = []
+with open("./requirements.txt", encoding="utf-8") as requirements_file:
+    # don't include peft yet until we check the int4
+    # need to manually install peft for now...
+    reqs = [r.strip() for r in requirements_file.readlines() if "peft" not in r]
+    reqs = [r for r in reqs if "flash-attn" not in r]
+    reqs = [r for r in reqs if r and r[0] != "#"]
+    for r in reqs:
+        install_requires.append(r)

 setup(
    name="axolotl",
-    version="0.3.0",
-    description="LLM Trainer",
-    long_description="Axolotl is a tool designed to streamline the fine-tuning of various AI models, offering support for multiple configurations and architectures.",
+    version="0.1",
+    description="You know you're going to axolotl questions",
    package_dir={"": "src"},
    packages=find_packages(),
    install_requires=install_requires,
-    dependency_links=dependency_links,
    extras_require={
+        "gptq": [
+            "alpaca_lora_4bit @ git+https://github.com/winglian/alpaca_lora_4bit.git@setup_pip",
+        ],
+        "gptq_triton": [
+            "alpaca_lora_4bit[triton] @ git+https://github.com/winglian/alpaca_lora_4bit.git@setup_pip",
+        ],
        "flash-attn": [
-            "flash-attn>=2.2.1",
+            "flash-attn==2.0.8",
        ],
        "extras": [
            "deepspeed",
        ],
+        "peft": [
+            "peft @ git+https://github.com/huggingface/peft.git",
+        ],
    },
 )
--- a/src/axolotl/logging_config.py
+++ b/src/axolotl/logging_config.py
@@ -23,7 +23,6 @@ class ColorfulFormatter(Formatter):
    }

    def format(self, record):
-        record.rank = int(os.getenv("LOCAL_RANK", "0"))
        log_message = super().format(record)
        return self.COLORS.get(record.levelname, "") + log_message + Fore.RESET

@@ -36,7 +35,7 @@ DEFAULT_LOGGING_CONFIG: Dict[str, Any] = {
        },
        "colorful": {
            "()": ColorfulFormatter,
-            "format": "[%(asctime)s] [%(levelname)s] [%(name)s.%(funcName)s:%(lineno)d] [PID:%(process)d] [RANK:%(rank)d] %(message)s",
+            "format": "[%(asctime)s] [%(levelname)s] [%(name)s.%(funcName)s:%(lineno)d] [PID:%(process)d] %(message)s",
        },
    },
    "filters": {},
--- a/src/axolotl/monkeypatch/llama_attn_hijack_flash.py
+++ b/src/axolotl/monkeypatch/llama_attn_hijack_flash.py
@@ -2,9 +2,7 @@

 # copied from https://github.com/lm-sys/FastChat/blob/main/fastchat/train/llama_flash_attn_monkey_patch.py

-import logging
 import warnings
-from functools import partial
 from typing import List, Optional, Tuple, Union

 import torch
@@ -35,9 +33,6 @@ except ImportError:
    )


-LOG = logging.getLogger("axolotl")
-
-
 def replace_llama_attn_with_flash_attn(packed: Optional[bool] = False):
    transformers.models.llama.modeling_llama.LlamaModel._prepare_decoder_attention_mask = (  # pylint: disable=protected-access
        _prepare_decoder_attention_mask
@@ -49,34 +44,6 @@ def replace_llama_attn_with_flash_attn(packed: Optional[bool] = False):
            llama_model_forward
        )

-    try:
-        from flash_attn.losses.cross_entropy import CrossEntropyLoss
-
-        LOG.info("patching with flash_attn.losses.cross_entropy")
-        transformers.models.llama.modeling_llama.CrossEntropyLoss = partial(
-            CrossEntropyLoss, inplace_backward=True
-        )
-    except ImportError:
-        LOG.info(
-            "optimized flash-attention CrossEntropyLoss not found (run `pip install 'git+https://github.com/Dao-AILab/flash-attention.git#egg=xentropy_cuda_lib&subdirectory=csrc/xentropy'`)"
-        )
-
-    try:
-        from flash_attn.ops.rms_norm import RMSNorm
-
-        class LlamaRMSNorm(RMSNorm):
-            """Patched LLamaRMSNorm"""
-
-            def __init__(self, hidden_size, eps=1e-6):
-                super().__init__(hidden_size, eps=eps)
-
-        LOG.info("patching with flash_attn.ops.rms_norm")
-        transformers.models.llama.modeling_llama.LlamaRMSNorm = LlamaRMSNorm
-    except ImportError:
-        LOG.info(
-            "optimized flash-attention RMSNorm not found (run `pip install 'git+https://github.com/Dao-AILab/flash-attention.git#egg=dropout_layer_norm&subdirectory=csrc/layer_norm'`)"
-        )
-

 # Disable the transformation of the attention mask in LlamaModel as the flash attention
 # requires the attention mask to be the same as the key_padding_mask
--- a/src/axolotl/prompters.py
+++ b/src/axolotl/prompters.py
@@ -309,6 +309,10 @@ class ShareGPTPrompter:  # pylint: disable=too-few-public-methods
        )

    def build_prompt(self, source) -> Generator[str, None, None]:
+        # ignore the system prompt if provided
+        if source[0]["from"] == "system":
+            source.pop(0)
+
        if len(source) < 2:
            # If there isn't a back and forth conversation, ignore it
            # also happens on the data splitting leaving empty conversations
@@ -317,12 +321,6 @@ class ShareGPTPrompter:  # pylint: disable=too-few-public-methods
            )

        conv = self._conversation.copy()
-
-        # Add the conversation system prompt if provided, otherwise use the default one
-        if source[0]["from"] == "system":
-            conv.system = source[0]["value"]
-            source.pop(0)
-
        roles = {"human": conv.roles[0], "gpt": conv.roles[1]}

        try:
--- a/src/axolotl/train.py
+++ b/src/axolotl/train.py
@@ -88,11 +88,6 @@ def train(
    if peft_config:
        LOG.info(f"Pre-saving adapter config to {cfg.output_dir}")
        peft_config.save_pretrained(cfg.output_dir)
-    # additionally presave the tokenizer and model configs
-    if not Path(cfg.output_dir).is_dir():
-        os.makedirs(cfg.output_dir, exist_ok=True)
-    tokenizer.save_pretrained(str(Path(cfg.output_dir)))
-    model.config.save_pretrained(str(Path(cfg.output_dir)))

    # In case we want to stop early with ctrl+c, this is a nice to have to save the pretrained model
    if cfg.local_rank == 0:
@@ -111,6 +106,9 @@ def train(
    if cfg.group_by_length:
        LOG.info("hang tight... sorting dataset for group_by_length")

+    if not Path(cfg.output_dir).is_dir():
+        os.makedirs(cfg.output_dir, exist_ok=True)
+    tokenizer.save_pretrained(cfg.output_dir)
    if cfg.flash_optimum:
        with torch.backends.cuda.sdp_kernel(
            enable_flash=True, enable_math=True, enable_mem_efficient=True
--- a/src/axolotl/utils/callbacks.py
+++ b/src/axolotl/utils/callbacks.py
@@ -11,6 +11,7 @@ import numpy as np
 import pandas as pd
 import torch
 import torch.distributed as dist
+from accelerate.state import PartialState
 from datasets import load_dataset
 from optimum.bettertransformer import BetterTransformer
 from tqdm import tqdm
@@ -24,13 +25,9 @@ from transformers.trainer_utils import PREFIX_CHECKPOINT_DIR, IntervalStrategy

 from axolotl.utils.bench import log_gpu_memory_usage
 from axolotl.utils.distributed import (
-    barrier,
-    broadcast_dict,
    gather_scalar_from_all_ranks,
    get_world_size,
-    is_distributed,
    is_main_process,
-    zero_first,
 )

 if TYPE_CHECKING:
@@ -38,6 +35,7 @@ if TYPE_CHECKING:

 LOG = logging.getLogger("axolotl.callbacks")
 IGNORE_INDEX = -100
+dist_state = PartialState()


 class SavePeftModelCallback(TrainerCallback):  # pylint: disable=too-few-public-methods
@@ -212,7 +210,7 @@ def bench_eval_callback_factory(trainer, tokenizer):
            "subject": example["subject"],
        }

-    with zero_first(is_main_process()):
+    with dist_state.main_process_first():
        bench_dataset = bench_dataset.map(tokenize_evals)
        bench_dataset = bench_dataset.filter(lambda x: x["labels"][-2] in abcd_idx)

@@ -260,7 +258,7 @@ def bench_eval_callback_factory(trainer, tokenizer):
            for s, p, r in zip(bench_name, preds, refs):  # pylint: disable=invalid-name
                bench_names[s]["preds"].append(p)
                bench_names[s]["refs"].append(r)
-            barrier()
+            dist_state.wait_for_everyone()
            local_bench_names = bench_names
            gathered_bench_names: List[Dict] = [{} for _ in range(get_world_size())]
            # Gather results from all GPUs to GPU 0
@@ -272,14 +270,10 @@ def bench_eval_callback_factory(trainer, tokenizer):
                lambda: len(data_loader), get_world_size()
            )

-            results = {}
-            if is_distributed() and not is_main_process():
+            if not is_main_process():
                dist.gather_object(local_bench_names, dst=0)
            else:
-                if is_distributed():
-                    dist.gather_object(local_bench_names, gathered_bench_names, dst=0)
-                else:
-                    gathered_bench_names = [local_bench_names]
+                dist.gather_object(local_bench_names, gathered_bench_names, dst=0)
                bench_loss = sum(loss_bench_ranks) / sum(len_data_loader_ranks)
                results = {f"{bench_split}_bench_loss": bench_loss}

@@ -318,8 +312,4 @@ def bench_eval_callback_factory(trainer, tokenizer):
                )["accuracy"]
                trainer.log(results)

-            results = broadcast_dict(results)
-            for key, val in results.items():
-                metrics[key] = val
-
    return BenchEvalCallback
--- a/src/axolotl/utils/config.py
+++ b/src/axolotl/utils/config.py
@@ -6,7 +6,6 @@ import os
 import torch

 from axolotl.utils.bench import log_gpu_memory_usage
-from axolotl.utils.models import load_model_config

 LOG = logging.getLogger("axolotl")

@@ -70,16 +69,6 @@ def normalize_config(cfg):
    else:
        cfg.torch_dtype = torch.float32

-    model_config = load_model_config(cfg)
-
-    # figure out if the model is llama
-    cfg.is_llama_derived_model = (
-        (hasattr(model_config, "model_type") and model_config.model_type == "llama")
-        or cfg.is_llama_derived_model
-        or "llama" in cfg.base_model
-        or (cfg.model_type and "llama" in cfg.model_type.lower())
-    )
-
    log_gpu_memory_usage(LOG, "baseline", cfg.device)


@@ -97,11 +86,6 @@ def validate_config(cfg):
            )
        )

-    if cfg.sample_packing and not cfg.pad_to_sequence_len:
-        LOG.warning(
-            "`pad_to_sequence_len: true` is recommended when using sample_packing"
-        )
-
    if cfg.gradient_accumulation_steps and cfg.batch_size:
        raise ValueError(
            "please set only one of gradient_accumulation_steps or batch_size"
@@ -113,7 +97,9 @@ def validate_config(cfg):
            "To calculate the equivalent gradient_accumulation_steps, divide batch_size / micro_batch_size / number of gpus.",
        )
    if cfg.load_4bit:
-        raise ValueError("cfg.load_4bit parameter has been deprecated")
+        raise ValueError(
+            "cfg.load_4bit parameter has been deprecated and replaced by cfg.gptq"
+        )

    if cfg.adapter == "qlora":
        if cfg.merge_lora:
@@ -220,15 +206,6 @@ def validate_config(cfg):
            "sample_packing not compatible with xformers_attention. Use flash_attention"
        )

-    if cfg.early_stopping_patience:
-        if not cfg.save_steps or not cfg.eval_steps:
-            raise ValueError(
-                "`early_stopping_patience` requires save_steps and eval_steps to be set. eval_steps should evenly divide save_steps."
-            )
-        if cfg.save_steps % cfg.eval_steps != 0:
-            raise ValueError(
-                "`early_stopping_patience` requires that eval_steps should evenly divide save_steps."
-            )
    # TODO
    # MPT 7b
    # https://github.com/facebookresearch/bitsandbytes/issues/25
--- a/src/axolotl/utils/data.py
+++ b/src/axolotl/utils/data.py
@@ -2,10 +2,12 @@
 import functools
 import hashlib
 import logging
+from hashlib import md5
 from pathlib import Path
 from typing import Tuple, Union

 import torch
+from accelerate.state import PartialState
 from datasets import (
    Dataset,
    DatasetDict,
@@ -41,7 +43,6 @@ from axolotl.prompters import (
    SummarizeTLDRPrompter,
 )
 from axolotl.utils.dict import DictDefault
-from axolotl.utils.distributed import is_main_process, zero_first
 from axolotl.utils.trainer import (
    calculate_total_num_steps,
    process_datasets_for_packing,
@@ -49,18 +50,12 @@ from axolotl.utils.trainer import (

 LOG = logging.getLogger("axolotl")
 DEFAULT_DATASET_PREPARED_PATH = "last_run_prepared"
-
-
-def md5(to_hash: str, encoding: str = "utf-8") -> str:
-    try:
-        return hashlib.md5(to_hash.encode(encoding), usedforsecurity=False).hexdigest()
-    except TypeError:
-        return hashlib.md5(to_hash.encode(encoding)).hexdigest()  # nosec
+state = PartialState()


 def prepare_dataset(cfg, tokenizer):
    if not cfg.pretraining_dataset:
-        with zero_first(is_main_process()):
+        with state.main_process_first():
            train_dataset, eval_dataset = load_prepare_datasets(
                tokenizer, cfg, DEFAULT_DATASET_PREPARED_PATH
            )
@@ -75,7 +70,7 @@ def prepare_dataset(cfg, tokenizer):
        train_dataset = train_dataset.with_format("torch")
        eval_dataset = None

-    with zero_first(is_main_process()):
+    with state.main_process_first():
        train_dataset, eval_dataset = process_datasets_for_packing(
            cfg, train_dataset, eval_dataset
        )
@@ -94,7 +89,7 @@ def load_tokenized_prepared_datasets(
 ) -> DatasetDict:
    tokenizer_name = tokenizer.__class__.__name__
    ds_hash = str(
-        md5(
+        md5(  # nosec
            (
                str(cfg.sequence_len)
                + "@"
@@ -103,8 +98,8 @@ def load_tokenized_prepared_datasets(
                )
                + "|"
                + tokenizer_name
-            )
-        )
+            ).encode("utf-8")
+        ).hexdigest()
    )
    prepared_ds_path = (
        Path(cfg.dataset_prepared_path) / ds_hash
@@ -380,7 +375,7 @@ def load_prepare_datasets(
        # see if we can go ahead and load the stacked dataset
        seed = f"@{str(cfg.seed)}" if cfg.seed else ""
        ds_hash = str(
-            md5(
+            md5(  # nosec
                (
                    str(cfg.sequence_len)
                    + "@"
@@ -391,8 +386,8 @@ def load_prepare_datasets(
                    )
                    + "|"
                    + tokenizer_name
-                )
-            )
+                ).encode("utf-8")
+            ).hexdigest()
        )
        prepared_ds_path = (
            Path(cfg.dataset_prepared_path) / ds_hash
@@ -506,10 +501,14 @@ def load_prepare_datasets(
            + "|"
            + str(cfg.seed or 42)
        )
-        train_fingerprint = md5(to_hash_train)
-        test_fingerprint = md5(to_hash_test)
+        train_fingerprint = hashlib.md5(
+            to_hash_train.encode(), usedforsecurity=False
+        ).hexdigest()
+        test_fingerprint = hashlib.md5(
+            to_hash_test.encode(), usedforsecurity=False
+        ).hexdigest()

-        with zero_first(is_main_process()):
+        with state.main_process_first():
            dataset = dataset.train_test_split(
                test_size=cfg.val_set_size,
                shuffle=False,
--- a/src/axolotl/utils/distributed.py
+++ b/src/axolotl/utils/distributed.py
@@ -1,30 +1,27 @@
 """
 utility helpers for distributed checks
 """
-import os
-import pickle  # nosec
-from contextlib import contextmanager
-
 import torch
 import torch.distributed as dist
-from accelerate import Accelerator
+from accelerate import DistributedType
+from accelerate.state import PartialState
+from accelerate.utils import wait_for_everyone

 accelerate = None  # pylint: disable=invalid-name

-
-def load_accelerate():
-    global accelerate  # pylint: disable=global-statement
-    accelerate = Accelerator()
+state = PartialState()


 def is_distributed():
    """
    Check if distributed training is initialized.
    """
-    global accelerate  # pylint: disable=global-statement
-    if not accelerate:
-        accelerate = Accelerator()
-    return dist.is_available() and dist.is_initialized()
+    return state.distributed_type in (
+        DistributedType.MULTI_GPU,
+        DistributedType.MULTI_CPU,
+        DistributedType.DEEPSPEED,
+        DistributedType.FSDP,
+    )


 def barrier():
@@ -32,34 +29,19 @@ def barrier():
    Acts as a barrier to wait for all processes. This ensures that all processes
    reach the barrier before proceeding further.
    """
-    if is_distributed():
-        dist.barrier()
+    wait_for_everyone()


-def is_main_process():
+def is_main_process() -> bool:
    """
    Check if the current process is the main process.
    If not in distributed mode, always return True.
    """
-    if not is_distributed():
-        return True
-    return dist.get_rank() == 0
+    return state.is_main_process


-def get_world_size():
-    return int(os.getenv("WORLD_SIZE", "1"))
-
-
-@contextmanager
-def zero_first(is_main):
-    """
-    runs the wrapped context so that rank 0 runs first before other ranks
-    """
-    if not is_main:  # other ranks wait first
-        barrier()
-    yield
-    if is_main:  # then rank 0 waits after it has run the context
-        barrier()
+def get_world_size() -> int:
+    return state.num_processes


 def gather_scalar_from_all_ranks(fn, world_size=1):  # pylint: disable=invalid-name
@@ -75,11 +57,9 @@ def gather_scalar_from_all_ranks(fn, world_size=1):  # pylint: disable=invalid-n
    - A list of computed values from all ranks if on the gathering rank, otherwise None.
    """
    value_scalar = fn()
-    if not is_distributed():
-        return [value_scalar]
    value_tensor = torch.tensor(value_scalar, device=dist.get_rank()).float()

-    if not is_main_process():
+    if not state.is_main_process:
        dist.gather(value_tensor, dst=0)
    else:
        gathered_tensors = [torch.zeros_like(value_tensor) for _ in range(world_size)]
@@ -94,30 +74,3 @@ def gather_scalar_from_all_ranks(fn, world_size=1):  # pylint: disable=invalid-n
                gathered_values.append(float(tensor.item()))
        return gathered_values
    return None
-
-
-def broadcast_dict(vals: dict):
-    if not is_distributed():
-        return vals
-
-    if is_main_process():
-        data_byte = pickle.dumps(vals)
-        data_tensor = torch.ByteTensor(list(data_byte)).to("cuda")
-        data_size = torch.IntTensor([len(data_byte)]).to("cuda")
-    else:
-        data_tensor = torch.empty([1024], dtype=torch.uint8, device="cuda")
-        data_size = torch.IntTensor([0]).to("cuda")
-
-    dist.broadcast(data_size, 0)
-    if not is_main_process():
-        # resize
-        data_tensor = data_tensor.new_empty([data_size.item()])
-
-    dist.broadcast(data_tensor, 0)
-
-    if not is_main_process():
-        data_list = data_tensor.cpu().tolist()
-        data_byte = bytes(data_list[: data_size.item()])
-        vals = pickle.loads(data_byte)  # nosec
-
-    return vals
--- a/src/axolotl/utils/models.py
+++ b/src/axolotl/utils/models.py
@@ -4,19 +4,19 @@
 import logging
 import math
 import os
+from pathlib import Path
 from typing import Optional, Tuple  # noqa: F401

 import bitsandbytes as bnb
 import torch
 import transformers
 from optimum.bettertransformer import BetterTransformer
-from peft import PeftConfig, prepare_model_for_kbit_training
+from peft import PeftConfig
 from transformers import (  # noqa: F401
    AutoConfig,
    AutoModelForCausalLM,
    AutoTokenizer,
    BitsAndBytesConfig,
-    GPTQConfig,
    LlamaConfig,
    PreTrainedModel,
    PreTrainedTokenizerBase,
@@ -155,17 +155,32 @@ def load_model(
        LOG.info("patching _expand_mask")
        hijack_expand_mask()

+    try:
+        if cfg.gptq:
+            from alpaca_lora_4bit.monkeypatch.peft_tuners_lora_monkey_patch import (
+                replace_peft_model_with_int4_lora_model,
+            )
+
+            replace_peft_model_with_int4_lora_model()
+    except Exception as err:
+        LOG.exception(err)
+        raise err
+
+    if not cfg.gptq and (
+        (cfg.adapter == "lora" and load_in_8bit)
+        or (cfg.adapter == "qlora" and cfg.load_in_4bit)
+    ):
+        try:
+            from peft import prepare_model_for_kbit_training
+        except ImportError:
+            # For backward compatibility
+            from peft import (
+                prepare_model_for_int8_training as prepare_model_for_kbit_training,
+            )
+
    model_kwargs = {}
    if cfg.model_revision:
        model_kwargs["revision"] = cfg.model_revision
-    if cfg.gptq:
-        model_config = load_model_config(cfg)
-        if not hasattr(model_config, "quantization_config"):
-            LOG.warning("model config does not contain quantization_config information")
-        else:
-            model_kwargs["quantization_config"] = GPTQConfig(
-                **model_config.quantization_config
-            )
    if cfg.adapter == "qlora" and cfg.load_in_4bit:
        model_kwargs["quantization_config"] = BitsAndBytesConfig(
            load_in_4bit=True,
@@ -176,7 +191,45 @@ def load_model(
            bnb_4bit_quant_type="nf4",
        )
    try:
-        if cfg.is_llama_derived_model and not cfg.trust_remote_code and not cfg.gptq:
+        if cfg.gptq and cfg.is_llama_derived_model:
+            from alpaca_lora_4bit.autograd_4bit import load_llama_model_4bit_low_ram
+            from huggingface_hub import snapshot_download
+
+            try:
+                snapshot_download_kwargs = {}
+                if cfg.base_model_ignore_patterns:
+                    snapshot_download_kwargs[
+                        "ignore_patterns"
+                    ] = cfg.base_model_ignore_patterns
+                cache_model_path = Path(
+                    snapshot_download(base_model, **snapshot_download_kwargs)
+                )
+                files = (
+                    list(cache_model_path.glob("*.pt"))
+                    + list(cache_model_path.glob("*.safetensors"))
+                    + list(cache_model_path.glob("*.bin"))
+                )
+                if len(files) > 0:
+                    model_path = str(files[0])
+                else:
+                    LOG.warning(
+                        "unable to find a cached model file, this will likely fail..."
+                    )
+                    model_path = str(cache_model_path)
+            except Exception:  # pylint: disable=broad-exception-caught
+                model_path = cfg.base_model
+            model, _ = load_llama_model_4bit_low_ram(
+                base_model_config if base_model_config else base_model,
+                model_path,
+                device_map=cfg.device_map,
+                half=cfg.fp16,
+                groupsize=cfg.gptq_groupsize if cfg.gptq_groupsize else -1,
+                is_v1_model=cfg.gptq_model_v1
+                if cfg.gptq_model_v1 is not None
+                else True,
+            )
+            load_in_8bit = False
+        elif cfg.is_llama_derived_model and not cfg.trust_remote_code:
            from transformers import LlamaForCausalLM

            config_kwargs = {}
@@ -222,24 +275,15 @@ def load_model(
        #     )
        #     model.train() # sets to train instead of eval mode
        elif model_type and not cfg.trust_remote_code:
-            if cfg.gptq:
-                model = AutoModelForCausalLM.from_pretrained(
-                    base_model,
-                    device_map=cfg.device_map,
-                    torch_dtype=cfg.torch_dtype,
-                    trust_remote_code=cfg.trust_remote_code or False,
-                    **model_kwargs,
-                )
-            else:
-                model = getattr(transformers, model_type).from_pretrained(
-                    base_model,
-                    device_map=cfg.device_map,
-                    load_in_8bit=cfg.load_in_8bit and cfg.adapter is not None,
-                    load_in_4bit=cfg.load_in_4bit and cfg.adapter is not None,
-                    torch_dtype=cfg.torch_dtype,
-                    trust_remote_code=cfg.trust_remote_code or False,
-                    **model_kwargs,
-                )
+            model = getattr(transformers, model_type).from_pretrained(
+                base_model,
+                device_map=cfg.device_map,
+                load_in_8bit=cfg.load_in_8bit and cfg.adapter is not None,
+                load_in_4bit=cfg.load_in_4bit and cfg.adapter is not None,
+                torch_dtype=cfg.torch_dtype,
+                trust_remote_code=cfg.trust_remote_code or False,
+                **model_kwargs,
+            )
        else:
            config = AutoConfig.from_pretrained(
                base_model,
@@ -315,12 +359,11 @@ def load_model(
                module.to(torch.float32)

    needs_fa2_dtype = cfg.adapter or cfg.fsdp
-    if (cfg.adapter == "lora" and load_in_8bit) or (
-        cfg.adapter == "qlora" and cfg.load_in_4bit
+    if not cfg.gptq and (
+        (cfg.adapter == "lora" and load_in_8bit)
+        or (cfg.adapter == "qlora" and cfg.load_in_4bit)
    ):
        LOG.info("converting PEFT model w/ prepare_model_for_kbit_training")
-        if cfg.gradient_checkpointing:
-            model.gradient_checkpointing_enable()
        model = prepare_model_for_kbit_training(
            model, use_gradient_checkpointing=cfg.gradient_checkpointing
        )
@@ -342,10 +385,22 @@ def load_model(
    if cfg.ddp and not load_in_8bit:
        model.to(f"cuda:{cfg.local_rank}")

+    if cfg.gptq:
+        # Scales to half
+        LOG.info("Fitting 4bit scales and zeros to half")
+        for _, module in model.named_modules():
+            if "Autograd4bitQuantLinear" in str(type(module)) or "Linear4bitLt" in str(
+                type(module)
+            ):
+                if hasattr(module, "is_v1_model") and module.is_v1_model:
+                    module.zeros = module.zeros.half()
+                module.scales = module.scales.half()
+                module.bias = module.bias.half()
+
    if (
        torch.cuda.device_count() > 1
        and int(os.getenv("WORLD_SIZE", "1")) > 1
-        and (cfg.load_in_4bit)
+        and (cfg.gptq or cfg.load_in_4bit)
    ):
        # llama is PROBABLY model parallelizable, but the default isn't that it is
        # so let's only set it for the 4bit, see
--- a/src/axolotl/utils/trainer.py
+++ b/src/axolotl/utils/trainer.py
@@ -33,7 +33,6 @@ from axolotl.utils.callbacks import (
 )
 from axolotl.utils.collators import DataCollatorForSeq2Seq
 from axolotl.utils.dataloader import MultipackDistributedDataloader
-from axolotl.utils.distributed import is_main_process, zero_first
 from axolotl.utils.schedulers import get_cosine_schedule_with_quadratic_warmup

 LOG = logging.getLogger("axolotl")
@@ -376,17 +375,14 @@ def disable_datasets_caching():

 def process_datasets_for_packing(cfg, train_dataset, eval_dataset):
    drop_long = partial(drop_long_seq, sequence_len=cfg.sequence_len)
-    with zero_first(is_main_process()):
-        train_dataset = train_dataset.filter(drop_long, num_proc=os.cpu_count())
-        if eval_dataset:
-            eval_dataset = eval_dataset.filter(drop_long, num_proc=os.cpu_count())
+    train_dataset = train_dataset.filter(drop_long, num_proc=os.cpu_count())
+    if eval_dataset:
+        eval_dataset = eval_dataset.filter(drop_long, num_proc=os.cpu_count())

-        if cfg.sample_packing:
-            train_dataset = train_dataset.map(add_position_ids, num_proc=os.cpu_count())
-            if eval_dataset:
-                eval_dataset = eval_dataset.map(
-                    add_position_ids, num_proc=os.cpu_count()
-                )
+    if cfg.sample_packing:
+        train_dataset = train_dataset.map(add_position_ids, num_proc=os.cpu_count())
+        if eval_dataset:
+            eval_dataset = eval_dataset.map(add_position_ids, num_proc=os.cpu_count())
    return train_dataset, eval_dataset


@@ -518,7 +514,23 @@ def setup_trainer(cfg, train_dataset, eval_dataset, model, tokenizer, total_num_
        training_arguments_kwargs["seed"] = cfg.seed

    if cfg.gradient_checkpointing:
-        training_arguments_kwargs["gradient_checkpointing"] = cfg.gradient_checkpointing
+        if cfg.gptq:
+            from alpaca_lora_4bit.gradient_checkpointing import (
+                apply_gradient_checkpointing,
+            )
+
+            gradient_checkpointing_ratio = (
+                cfg.gradient_checkpointing_ratio
+                if cfg.gradient_checkpointing_ratio
+                else 1.0
+            )
+            apply_gradient_checkpointing(
+                model, checkpoint_ratio=gradient_checkpointing_ratio
+            )
+        else:
+            training_arguments_kwargs[
+                "gradient_checkpointing"
+            ] = cfg.gradient_checkpointing
    if cfg.fsdp:
        training_arguments_kwargs["fsdp"] = cfg.fsdp
        if cfg.fsdp_config:
@@ -576,10 +588,6 @@ def setup_trainer(cfg, train_dataset, eval_dataset, model, tokenizer, total_num_
        training_arguments_kwargs["do_bench_eval"] = cfg.do_bench_eval
        if cfg.bench_dataset:
            training_arguments_kwargs["bench_dataset"] = cfg.bench_dataset
-    if cfg.metric_for_best_model:
-        training_arguments_kwargs["metric_for_best_model"] = cfg.metric_for_best_model
-    if cfg.greater_is_better:
-        training_arguments_kwargs["greater_is_better"] = cfg.greater_is_better

    # DDP Config
    if cfg.ddp_timeout:
@@ -605,10 +613,11 @@ def setup_trainer(cfg, train_dataset, eval_dataset, model, tokenizer, total_num_
        output_dir=cfg.output_dir,
        save_total_limit=cfg.save_total_limit if cfg.save_total_limit else 4,
        load_best_model_at_end=(
-            (cfg.load_best_model_at_end is not False or cfg.early_stopping_patience)
+            cfg.load_best_model_at_end is not False
            and cfg.val_set_size > 0
            and cfg.save_steps
            and cfg.save_steps % cfg.eval_steps == 0
+            and cfg.load_in_8bit is not True
        )
        or False,
        ddp_find_unused_parameters=False if cfg.ddp else None,
@@ -640,6 +649,13 @@ def setup_trainer(cfg, train_dataset, eval_dataset, model, tokenizer, total_num_
    if cfg.relora_steps:
        callbacks.append(ReLoRACallback(cfg))

+    # TODO on_save callback to sync checkpoints to GCP/AWS in background
+    if cfg.early_stopping_patience:
+        early_stop_cb = EarlyStoppingCallback(
+            cfg.early_stopping_patience,
+        )
+        callbacks.append(early_stop_cb)
+
    if cfg.local_rank == 0 and cfg.adapter in [
        "lora",
        "qlora",
@@ -706,11 +722,4 @@ def setup_trainer(cfg, train_dataset, eval_dataset, model, tokenizer, total_num_
    if cfg.do_bench_eval:
        trainer.add_callback(bench_eval_callback_factory(trainer, tokenizer))

-    # TODO on_save callback to sync checkpoints to GCP/AWS in background
-    if cfg.early_stopping_patience:
-        early_stop_cb = EarlyStoppingCallback(
-            cfg.early_stopping_patience,
-        )
-        trainer.add_callback(early_stop_cb)
-
    return trainer
--- a/tests/test_data.py
+++ b/tests/test_data.py
@@ -1,64 +0,0 @@
-"""
-test module for the axolotl.utis.data module
-"""
-import unittest
-
-from transformers import LlamaTokenizer
-
-from axolotl.utils.data import encode_pretraining, md5
-
-
-class TestEncodePretraining(unittest.TestCase):
-    """
-    test class for encode pretraining and md5 helper
-    """
-
-    def setUp(self):
-        self.tokenizer = LlamaTokenizer.from_pretrained("huggyllama/llama-7b")
-        self.tokenizer.add_special_tokens(
-            {
-                "eos_token": "</s>",
-                "bos_token": "<s>",
-                "unk_token": "<unk>",
-                "pad_token": "<pad>",
-            }
-        )
-        self.max_tokens = 15  # set a small number for easy inspection
-
-    def test_encode_pretraining(self):
-        examples = {
-            "text": [
-                "Hello, world!",
-                "Nice to meet you.",
-                "lorem ipsum dolor sit amet.",
-                "Nice to meet you again!.",
-                "hello, hello",
-            ]
-        }
-        result = encode_pretraining(self.tokenizer, self.max_tokens, examples)
-
-        self.assertEqual(len(result["input_ids"]), 3)
-
-        # Assert the length of input_ids and attention_mask is correct
-        self.assertEqual(len(result["input_ids"][0]), self.max_tokens)
-        self.assertEqual(len(result["attention_mask"][0]), self.max_tokens)
-
-        # Assert EOS and PAD tokens are correctly added
-        # hello world! is 4 tokens
-        self.assertEqual(result["input_ids"][0][0], self.tokenizer.bos_token_id)
-        self.assertEqual(result["input_ids"][0][5], self.tokenizer.eos_token_id)
-        self.assertEqual(result["input_ids"][0][6], self.tokenizer.pad_token_id)
-        # second part, 5 tokens
-        self.assertEqual(result["input_ids"][0][7], self.tokenizer.bos_token_id)
-        self.assertEqual(result["input_ids"][0][13], self.tokenizer.eos_token_id)
-        self.assertEqual(result["input_ids"][0][14], self.tokenizer.pad_token_id)
-
-    def test_md5(self):
-        self.assertEqual(md5("hello world"), "5eb63bbbe01eeed093cb22bb8f5acdc3")
-        self.assertEqual(
-            md5("hello world", "utf-8"), "5eb63bbbe01eeed093cb22bb8f5acdc3"
-        )
-
-
-if __name__ == "__main__":
-    unittest.main()
--- a/tests/test_validation.py
+++ b/tests/test_validation.py
@@ -328,20 +328,6 @@ class ValidationTest(unittest.TestCase):
                for record in self._caplog.records
            )

-        cfg = DictDefault(
-            {
-                "sample_packing": True,
-                "pad_to_sequence_len": None,
-            }
-        )
-        with self._caplog.at_level(logging.WARNING):
-            validate_config(cfg)
-            assert any(
-                "`pad_to_sequence_len: true` is recommended when using sample_packing"
-                in record.message
-                for record in self._caplog.records
-            )
-
        cfg = DictDefault(
            {
                "max_packed_sequence_len": 2048,
Author	SHA1	Message	Date
Wing Lian	83d904a27d	fix the context manager call Some checks failed pre-commit / pre-commit (push) Has been cancelled Details PyTest / test (3.10) (push) Has been cancelled Details PyTest / test (3.9) (push) Has been cancelled Details	2023-09-03 22:49:28 -04:00
Wing Lian	5e4a760ad8	start to swap out for accelerate partial state	2023-09-03 22:41:29 -04:00