nit [skip-e2e]

nit
docs
2025-08-13 12:58:30 +01:00 · 2025-08-13 12:58:13 +01:00 · 2025-08-13 12:53:58 +01:00 · 2025-08-13 11:27:15 +01:00 · 2025-08-13 11:24:22 +01:00 · 2025-08-13 11:24:05 +01:00
14 changed files with 65 additions and 117 deletions
--- a/.github/CONTRIBUTING.md
+++ b/.github/CONTRIBUTING.md
@@ -57,6 +57,13 @@ We welcome ideas for improvements and new features. To suggest an enhancement, o
 5. Push your branch to your fork on GitHub.
 6. Open a new pull request against the `main` branch of the axolotl repository. Include a clear and concise description of your changes, referencing any related issues.

+#### Skipping CI Checks
+
+You can skip certain CI checks by including specific keywords in your commit messages:
+
+- `[skip ci]` or `skip ci` - Skips all CI checks for that commit
+- `[skip-e2e]` or `skip-e2e` - Skips only end-to-end tests while running other CI checks
+
 ## Style Guidelines

 ### Code Style
--- a/.github/workflows/main.yml
+++ b/.github/workflows/main.yml
@@ -98,6 +98,12 @@ jobs:
            python_version: "3.11"
            pytorch: 2.7.1
            axolotl_extras:
+            is_latest:
+          - cuda: 126
+            cuda_version: 12.6.3
+            python_version: "3.11"
+            pytorch: 2.7.1
+            axolotl_extras: vllm
            is_latest: true
          - cuda: 128
            cuda_version: 12.8.1
@@ -151,6 +157,18 @@ jobs:
            python_version: "3.11"
            pytorch: 2.6.0
            axolotl_extras:
+          - cuda: 126
+            cuda_version: 12.6.3
+            python_version: "3.11"
+            pytorch: 2.7.1
+            axolotl_extras:
+            is_latest:
+          - cuda: 126
+            cuda_version: 12.6.3
+            python_version: "3.11"
+            pytorch: 2.7.1
+            axolotl_extras: vllm
+            is_latest: true
    runs-on: axolotl-gpu-runner
    steps:
      - name: Checkout
--- a/.github/workflows/tests.yml
+++ b/.github/workflows/tests.yml
@@ -105,7 +105,8 @@ jobs:

      - name: Run tests
        run: |
-          pytest -v --durations=10 -n8 --dist loadfile --ignore=tests/e2e/ --ignore=tests/patched/ --ignore=tests/cli/ tests/ --cov=axolotl --cov-report=xml
+          pytest -v --durations=10 -n8 --dist loadfile --ignore=tests/e2e/ --ignore=tests/patched/ --ignore=tests/cli/ --ignore=tests/monkeypatch/ tests/ --cov=axolotl --cov-report=xml
+          pytest -v --durations=10 tests/monkeypatch/ --cov=axolotl --cov-append --cov-report=xml
          pytest -v --durations=10 tests/patched/ --cov=axolotl --cov-append --cov-report=xml
          pytest -v --durations=10 tests/cli/ --cov=axolotl --cov-append --cov-report=xml

@@ -179,8 +180,8 @@ jobs:

      - name: Run tests
        run: |
-          pytest -v --durations=10 -n8 --dist loadfile --ignore=tests/e2e/ --ignore=tests/patched/ --ignore=tests/cli/ tests/
-          pytest -v --durations=10 tests/patched/
+          pytest -v --durations=10 -n8 --dist loadfile --ignore=tests/e2e/ --ignore=tests/patched/ --ignore=tests/cli/ --ignore=tests/monkeypatch/ tests/ --cov=axolotl --cov-report=xml
+          pytest -v --durations=10 tests/monkeypatch/ --cov=axolotl --cov-append --cov-report=xml
          pytest -v --durations=10 tests/cli/

      - name: cleanup pip cache
--- a/.pre-commit-config.yaml
+++ b/.pre-commit-config.yaml
@@ -3,7 +3,7 @@ default_language_version:

 repos:
 -   repo: https://github.com/pre-commit/pre-commit-hooks
-    rev: v5.0.0
+    rev: v6.0.0
    hooks:
    -   id: check-yaml
    -   id: end-of-file-fixer
@@ -23,7 +23,7 @@ repos:
    hooks:
    - id: flake8
 -   repo: https://github.com/pylint-dev/pylint
-    rev: v3.3.7
+    rev: v3.3.8
    hooks:
    - id: pylint
 -   repo: https://github.com/pre-commit/mirrors-mypy
--- a/CITATION.cff
+++ b/CITATION.cff
@@ -0,0 +1,10 @@
+cff-version: 1.2.0
+type: software
+title: "Axolotl: Post-Training for AI Models"
+message: "If you use this software, please cite it as below."
+authors:
+  - name: "Axolotl maintainers and contributors"
+repository-code: "https://github.com/axolotl-ai-cloud/axolotl"
+url: "https://axolotl.ai/"
+license: Apache-2.0
+date-released: "2023-05-30"
--- a/README.md
+++ b/README.md
@@ -149,6 +149,20 @@ Contributions are welcome! Please see our [Contributing Guide](https://github.co

 Interested in sponsoring? Contact us at [wing@axolotl.ai](mailto:wing@axolotl.ai)

+## 📝 Citing Axolotl
+
+If you use Axolotl in your research or projects, please cite it as follows:
+
+```bibtex
+@software{axolotl,
+  title = {Axolotl: Post-Training for AI Models},
+  author = {{Axolotl maintainers and contributors}},
+  url = {https://github.com/axolotl-ai-cloud/axolotl},
+  license = {Apache-2.0},
+  year = {2023}
+}
+```
+
 ## 📜 License

 This project is licensed under the Apache 2.0 License - see the [LICENSE](LICENSE) file for details.
--- a/TODO.md
+++ b/TODO.md
@@ -1,10 +0,0 @@
-# todo list
-
- [] Validation of parameters for combinations that won't work
-
-
-
-## things that are known not to work
-
- FSDP offload and gradient_checkpointing - https://github.com/pytorch/pytorch/issues/82203
- adamw_bnb_8bit doesn't play well with FSDP offload
--- a/src/axolotl/init.py
+++ b/src/axolotl/init.py
@@ -4,4 +4,4 @@ import pkgutil

 __path__ = pkgutil.extend_path(__path__, __name__)  # Make this a namespace package

-__version__ = "0.12.1"
+__version__ = "0.13.0.dev"
--- a/src/axolotl/cli/config.py
+++ b/src/axolotl/cli/config.py
@@ -153,15 +153,14 @@ def prepare_plugins(cfg: DictDefault):
        plugin_manager = PluginManager.get_instance()
        for plugin_name in cfg["plugins"]:
            plugin_manager.register(plugin_name)
+        for plugin in plugin_manager.plugins.values():
+            plugin.register(cfg)


 def plugin_set_cfg(cfg: DictDefault):
    if cfg.get("plugins"):
        plugin_manager = PluginManager.get_instance()
        plugin_manager.cfg = cfg
-        # now that we have the finalized cfg, register the plugins individually
-        for plugin in plugin_manager.plugins.values():
-            plugin.register(cfg)


 def load_cfg(
--- a/src/axolotl/cli/utils/train.py
+++ b/src/axolotl/cli/utils/train.py
@@ -67,14 +67,12 @@ def build_command(base_cmd: list[str], options: dict[str, Any]) -> list[str]:

 def generate_config_files(config: str, sweep: str | None) -> Iterator[tuple[str, bool]]:
    """
-    Generate list of configuration files to process.
+    Generate list of configuration files to process. Yields a tuple of the configuration file name and a boolean indicating
+    whether this is a group of configurations (i.e., a sweep).

    Args:
        config: Base configuration file
        sweep: Sweep configuration file
-
-    Yields:
-        Tuple of configuration file name and whether this is a group of configurations
    """

    if not sweep:
--- a/src/axolotl/integrations/base.py
+++ b/src/axolotl/integrations/base.py
@@ -76,8 +76,8 @@ class BasePlugin:
    def __init__(self):
        """Initializes the BasePlugin."""

-    def register(self, cfg: DictDefault):  # pylint: disable=unused-argument
-        """Registers the plugin with the given configuration.
+    def register(self, cfg: dict):  # pylint: disable=unused-argument
+        """Registers the plugin with the given configuration as an unparsed dict.

        Args:
            cfg: The configuration for the plugin.
--- a/src/axolotl/loaders/patch_manager.py
+++ b/src/axolotl/loaders/patch_manager.py
@@ -73,9 +73,6 @@ class PatchManager:
        self._apply_voxtral_patches()

    def _apply_transformers_patches(self):
-        from axolotl.monkeypatch.transformers.modeling_flash_attention_utils import (
-            patch_prepare_from_posids,
-        )
        from axolotl.monkeypatch.transformers.trainer_loss_calc import (
            patch_evaluation_loop,
            patch_maybe_log_save_evaluate,
@@ -87,7 +84,6 @@ class PatchManager:
            and self.cfg.fsdp_version == 2
        )

-        patch_prepare_from_posids()
        patch_evaluation_loop(patch_fsdp2)
        patch_maybe_log_save_evaluate()

--- a/src/axolotl/monkeypatch/transformers/modeling_flash_attention_utils.py
+++ b/src/axolotl/monkeypatch/transformers/modeling_flash_attention_utils.py
@@ -1,87 +0,0 @@
-"""
-Monkey patch to fix transformers.modeling_flash_attention_utils.
-
-see https://github.com/huggingface/transformers/pull/39653/files
-"""
-
-import sys
-
-import torch
-
-
-def _prepare_from_posids(query, key, value, position_ids):
-    """
-    This function returns necessary arguments to call `flash_attn_varlen_func`.
-    All three query, key, value states will be flattened.
-    Cumulative lengths of each examples in the batch will be extracted from position_ids.
-    NOTE: ideally cumulative lengths should be prepared at the data collator stage
-    Arguments:
-        query (`torch.Tensor`):
-            Query state with padding. Shape: (batch_size, query_length, num_heads, head_dim).
-        key (`torch.Tensor`):
-            Key state with padding. Shape: (batch_size, kv_seq_len, num_key_value_heads, head_dim).
-        value (`torch.Tensor`):
-            Value state with padding. Shape: (batch_size, kv_seq_len, num_key_value_heads, head_dim).
-        position_ids (`torch.Tensor`):
-            Boolean or int tensor of shape (batch_size, sequence_length), 1 means valid and 0 means not valid.
-    Return:
-        query (`torch.Tensor`):
-            Query state without padding. Shape: (total_target_length, num_heads, head_dim).
-        key (`torch.Tensor`):
-            Key state with padding. Shape: (total_source_length, num_key_value_heads, head_dim).
-        value (`torch.Tensor`):
-            Value state with padding. Shape: (total_source_length, num_key_value_heads, head_dim).
-        indices_q (`torch.Tensor`):
-            The indices of non-masked tokens from the flattened input target sequence.
-        (cu_seqlens_q, cu_seqlens_k) (`tuple[int]`):
-            The cumulative sequence lengths for the target (query) and source (key, value), used to index into ragged (unpadded) tensors. `cu_seqlens` shape is (batch_size + 1,).
-        (max_seqlen_in_batch_q, max_seqlen_in_batch_k) (`tuple[int]`):
-            Maximum sequence length in batch (`max_seqlen_in_batch_q` for the target sequence i.e. query, `max_seqlen_in_batch_k` for the source sequence i.e. key/value).
-    """
-    query = query.contiguous().view(-1, query.size(-2), query.size(-1))
-    key = key.contiguous().view(-1, key.size(-2), key.size(-1))
-    value = value.contiguous().view(-1, value.size(-2), value.size(-1))
-
-    position_ids = position_ids.flatten()
-    indices_q = torch.arange(
-        position_ids.size(0), device=position_ids.device, dtype=torch.int32
-    )
-
-    cu_seq_lens = torch.cat(
-        (
-            indices_q[position_ids == 0],
-            torch.tensor(
-                position_ids.size(), device=position_ids.device, dtype=torch.int32
-            ),
-        )
-    )
-    # NOTE: With torch compile, this will cause a graph break if you don't set
-    # `TORCHDYNAMO_CAPTURE_SCALAR_OUTPUTS=1` in the environment or call
-    # `torch._dynamo.config.capture_scalar_outputs = True` before doing the forward pass.
-    # This is a limitation of flash attention API, as the function `flash_attn_varlen_func`
-    # requires `max_length_q`, `max_length_k` to be passed as `int` and not `torch.Tensor`.
-    # https://github.com/Dao-AILab/flash-attention/blob/2dd8078adc1d9b74e315ee99718c0dea0de8eeb6/flash_attn/flash_attn_interface.py#L1423-L1424
-    # We should use cu_seq_lens instead of position_ids to get the max length since position_ids is not always increasing
-    # for some models (e.g. qwen2-vl).
-    max_length = cu_seq_lens.diff().max().item()
-    return (
-        query,
-        key,
-        value,
-        indices_q,
-        (cu_seq_lens, cu_seq_lens),
-        (max_length, max_length),
-    )
-
-
-def patch_prepare_from_posids():
-    import transformers.modeling_flash_attention_utils
-
-    transformers.modeling_flash_attention_utils._prepare_from_posids = (  # pylint: disable=protected-access
-        _prepare_from_posids
-    )
-    setattr(
-        sys.modules["transformers.modeling_flash_attention_utils"],
-        "_prepare_from_posids",
-        _prepare_from_posids,
-    )
--- a/tests/core/test_builders.py
+++ b/tests/core/test_builders.py
@@ -281,7 +281,9 @@ class TestHFRLTrainerBuilder:
        # Other settings
        assert training_arguments.dataloader_num_workers == 1
        assert training_arguments.dataloader_pin_memory is True
-        assert training_arguments.gradient_checkpointing is False
+
+        # TODO(wing): restore once trl releases 0.22.0
+        # assert training_arguments.gradient_checkpointing is True

    def test_dpo_training_arguments(self, dpo_cfg, model, tokenizer):
        builder = HFRLTrainerBuilder(dpo_cfg, model, tokenizer)
Author	SHA1	Message	Date
Salman Mohammadi	7d8e8c9ac2	nit [skip-e2e]	2025-08-13 12:58:30 +01:00
Salman Mohammadi	7c2466b739	nit	2025-08-13 12:58:13 +01:00
Salman Mohammadi	3146cb56dd	docs	2025-08-13 12:53:58 +01:00
Salman Mohammadi	c09b0a3bbf	reverting change	2025-08-13 11:27:15 +01:00
Salman Mohammadi	e05acccd77	linting	2025-08-13 11:24:22 +01:00
Salman Mohammadi	c44abad531	debugging CI	2025-08-13 11:24:05 +01:00
Salman Mohammadi	817d70e669	debugging CI	2025-08-13 10:45:41 +01:00
Salman Mohammadi	03f5a7fd16	adding back	2025-08-12 18:35:12 +01:00
Salman Mohammadi	3d9b96a94f	testing revert	2025-08-12 15:53:43 +01:00
Salman Mohammadi	42c16024a2	docs	2025-08-12 15:34:46 +01:00
Salman Mohammadi	ec94d632f3	docs	2025-08-12 14:07:55 +01:00
Salman Mohammadi	e8bd3b0b3b	Merge branch 'fix-preview' of github.com:axolotl-ai-cloud/axolotl into fix-preview	2025-08-12 13:42:56 +01:00
Salman Mohammadi	5a08b94668	update workflow	2025-08-12 12:29:09 +01:00
salman	ecb8c1f4b3	Merge branch 'main' into fix-preview	2025-08-12 09:43:39 +01:00
Salman Mohammadi	ab57be6526	render docs on python file change to preview api ref	2025-08-12 09:43:23 +01:00
Wing Lian	3d45620008	remove prepare-from-posids patch (#3052 ) [skip ci]	2025-08-11 09:34:41 -04:00
github-actions[bot]	ce20e838b5	chore: update pre-commit hooks (#3050 ) [skip ci] Co-authored-by: djsaunde <1245942+djsaunde@users.noreply.github.com>	2025-08-11 09:32:21 -04:00
Wing Lian	d4d84d48af	fix ray train and add fsdp2 smoke test for ray trainer (#3053 ) * add fsdp2 smokle test for ray trainer * fix raytrain with fsdp2	2025-08-11 09:31:54 -04:00
Wing Lian	c9640bca2c	attempt to fix quartodoc render for yields	2025-08-10 22:23:09 -04:00
Wing Lian	9b12c05660	use exec instead of subprocess to make ctrl+c nicer for cli (#3044 ) * use exec instead of subprocess to make ctrl+c nicer for cli * change var name to use_exec * simplify to bool * flush std* * patch subprocess as mock in test * fix tests * more test fixes	2025-08-10 20:22:20 -04:00
Wing Lian	686933194e	fix vllm tagging and add cloud images w/o tmux (#3049 ) [skip ci]	2025-08-10 20:21:56 -04:00
Wing Lian	d12b461d19	follow up fix for plugin registration (#3054 ) [skip ci]	2025-08-10 20:21:38 -04:00
Wing Lian	d6b81b3683	update training args check for new defaults (#3051 ) [skip ci] * update training args check for new defaults * skip check for now	2025-08-10 11:26:22 -04:00
Wing Lian	05f1b4b2e8	run monkeypatch tests in seperate runner (#3047 )	2025-08-09 14:34:07 -04:00
Wing Lian	7cfc80ec77	set dev version (#3045 ) [skip ci]	2025-08-08 13:56:53 -04:00
salman	0da6a95efa	Add citation.tff (#3043 ) [skip ci]	2025-08-08 16:18:42 +01:00