axolotl

Author	SHA1	Message	Date
Wing Lian	29cf15a28c	improve save callbacks (#1592 )	2024-05-04 23:19:18 -04:00
Chirag Jain	dde02fcb94	Pass weakref to model in the SIGINT handler to free up model post train function (#1581 ) * Pass weakref to model in the SIGINT handler to free up model post train() * Fix lint issues * chore: lint --------- Co-authored-by: Wing Lian <wing.lian@gmail.com>	2024-05-03 11:05:28 -04:00
Ali Mosavian	b9bb169602	FIX: TRL trainer preprocessing step was running in one process (#1583 ) * FIX: TRL trainer preprocessing step was running in one process * FIX: Changed so that dataset_num_proc is sent to CPO, KTO and ORPO trainer args and directly to the trainer when DPO * FIX: Changed back to only support ORPO for now, since KTO is handled in another way --------- Co-authored-by: Ali Mosavian <ali.mosavian@kry.se>	2024-05-03 11:02:59 -04:00
JohanWork	601c08b4c2	ADD: warning hub model (#1301 ) * update warning for save_strategy * update * clean up * update * Update test_validation.py * fix validation step * update * test_validation * update * fix * fix --------- Co-authored-by: NanoCode012 <kevinvong@rocketmail.com>	2024-05-01 01:05:12 +09:00
Abhinand	cc5d31e0d9	Add debug option for RL dataset preprocessing (#1404 ) * adding debug option for RL dataset preprocessing * Refine formatting of debugging code in RL dataset preprocessing * Update __init__.py * chore: fix lint --------- Co-authored-by: NanoCode012 <kevinvong@rocketmail.com>	2024-05-01 00:36:04 +09:00
NanoCode012	1aeece6e24	chore(doc): clarify micro_batch_size (#1579 ) [skip ci]	2024-05-01 00:33:53 +09:00
Wing Lian	5294653a2d	PoSE context length ext (#1567 ) * PoSE wip * fixes for pose splitting * set pose context len so we can pick that up seperately from the usable training context len * support min sample len and define num chunks * fix chunk splitting * support for curriculum/ordered learning with pose * fix sequence len sort * add curriculum_sampling to pydantic	2024-04-27 12:28:20 -04:00
Motoki Wu	98c25e15cb	Add ORPO example and e2e test (#1572 ) * add example for mistral orpo * sample_packing: false for orpo * go to load_dataset (since load_rl_datasets require a transfom_fn, which only dpo uses currently)	2024-04-27 12:07:06 -04:00
Wing Lian	68601ec6ad	make sure everything stays in the same dtype when using dpo + FSDP (#1559 )	2024-04-22 16:00:05 -04:00
Haoxiang Wang	60f5ce0569	Add support for Gemma chat template (#1530 ) * Add support for Gemma chat template * Update fschat version to include its newest support for Gemma chat style * pin fastchat to current HEAD --------- Co-authored-by: Wing Lian <wing.lian@gmail.com>	2024-04-21 19:55:40 -04:00
Frank Ruis	7477a53287	wrap prepared_ds_path in str() to avoid TypeError in fsspec package (#1548 ) * wrap prepared_ds_path in str() to avoid TypeError in fsspec package `fsspec` calls `if "::" in path` on `prepared_ds_path`, which will throw an error if it is a `PosixPath` object. * update test too --------- Co-authored-by: Wing Lian <wing.lian@gmail.com>	2024-04-21 19:55:20 -04:00
Wing Lian	7d1d22f72f	ORPO Trainer replacement (#1551 ) * WIP use trl ORPOTrainer * fixes to make orpo work with trl * fix the chat template laoding * make sure to handle the special tokens and add_generation for assistant turn too	2024-04-19 17:25:36 -04:00
NanoCode012	0e8f340945	fix(yml): update llama-3 config (#1543 ) [skip ci]	2024-04-19 20:44:46 +09:00
NanoCode012	59ef25470c	fix(packages): lock datasets version (#1545 )	2024-04-19 20:42:10 +09:00
Wing Lian	c10563c444	fix broken linting (#1541 ) * chore: lint * include examples in yaml check * mistral decided to gate their models... * more mistral models that were gated	2024-04-19 01:03:04 -04:00
Monk (looking for PhD Fall’24)	37c037c69d	Adding Llama-3 qlora (#1536 ) * Create qlora.yml * Update qlora.yml	2024-04-18 21:27:32 +02:00
Wing Lian	15f7910d33	llama-3 examples (#1537 )	2024-04-18 14:28:03 -04:00
NanoCode012	d28ba2e405	feat(doc): Add example for pad_token (#1535 )	2024-04-19 02:20:20 +09:00
Atlas	0eadfc8c86	Create mixtral_22.yml (#1514 ) [skip ci] Code sourced from here: https://twitter.com/mattshumer_/status/1778135774887567712	2024-04-17 01:16:00 -04:00
Atlas	bcaa92325d	Update Readme to include support for Mixtral8X22B (#1518 ) [skip ci]	2024-04-17 01:15:30 -04:00
YTING	7d9bafcb88	Update README.md (#1521 ) [skip ci]	2024-04-17 01:15:05 -04:00
Wing Lian	e07dcb288c	add docs around pre-processing (#1529 )	2024-04-16 19:45:46 -04:00
Wing Lian	6319da1f9b	Unsloth gradient checkpointing offload (#1528 ) * unsloth gradient checkpointing * fix validation too * fixes to make it work with mistral * monkeypatch the checkpoint fn earlier	2024-04-16 14:53:57 -04:00
Wing Lian	132eb740f0	DBRX Model Support (#1462 ) * wip for dbrx finetuning * add fastcore for parallel loading of sharded weights * fix dtype for load, use PartialState instead of accelerator to init process group, remove redundant wandb callback * update to use v2 of the converted model * more fixes for dbrx loras * make sure to enable fsdp activation checkpointing * fix support for 8bit loras too for dbrx * apply z3 leaf moe fix for DBRX with deepspeed * don't raise value error since child module searches could fail and be ok * revert a previous change to fix fsdp * update mistral/mistral qlora+fsdp yamls * fix qlora+fsdp quant storage type * more edge cases for qlora-fsdp * fixes for fsdp+qlora w optimizer in 8bit * add bigstral z3 config and make sure to use full_state_dict for fsdp	2024-04-12 09:02:36 -04:00
Thomas Capelle	5ed29393e3	Update SaveAxolotlConfigtoWandBCallback to use artifact instead of save (#1483 ) * deprecated wandb.save * also use wandb.save for axolotl yaml * chore: lint --------- Co-authored-by: Wing Lian <wing.lian@gmail.com>	2024-04-09 18:58:38 -04:00
Wing Lian	da9b1a3196	use locale agnostic seperator to make large nums easier to read (#1503 )	2024-04-09 17:28:43 -04:00
DavidFarago	057fa44191	WIP: Support table logging for mlflow, too (#1506 ) * WIP: Support table logging for mlflow, too Create a `LogPredictionCallback` for both "wandb" and "mlflow" if specified. In `log_prediction_callback_factory`, create a generic table and make it specific only if the newly added `logger` argument is set to "wandb" resp. "mlflow". See https://github.com/OpenAccess-AI-Collective/axolotl/issues/1505 * chore: lint * add additional clause for mlflow as it's optional * Fix circular imports --------- Co-authored-by: Dave Farago <dfarago@innoopract.com> Co-authored-by: Wing Lian <wing.lian@gmail.com>	2024-04-09 17:28:27 -04:00
Scott Fleming	8fa0785f74	Correctly handle splits for datasets.arrow_dataset.Dataset objects (#1504 ) * Correctly handle splits for datasets.arrow_dataset.Dataset objects The `load_tokenized_prepared_datasets` function currently has logic for loading a dataset from local path that always checks if a split is in the dataset. The problem is, if the dataset is loaded using `load_from_disk` and it is an Arrow-based dataset, there is no split information. Instead what happens is, by calling `split in ds`, it presumably searches through all the rows and columns of the arrow dataset object to find e.g., 'train' assuming `split == 'train'`. This causes the program to hang. See https://chat.openai.com/share/0d567dbd-d60b-4079-9040-e1de58a4dff3 for context. * chore: lint --------- Co-authored-by: Wing Lian <wing.lian@gmail.com>	2024-04-09 16:40:26 -04:00
Wing Lian	4313b1a6a0	Print versions (#1496 ) * print out dependency versions for easier debugging * improve readability	2024-04-09 11:05:15 -04:00
Maziyar Panahi	7f17eff81a	Fix the wrong adapter in qwen2-moe-qlora example (#1501 ) [skip ci] It should be `qlora` instead of `lora`	2024-04-09 10:57:24 -04:00
Wing Lian	ff01c45127	add field to sft dataset pydantic for completion support (#1497 )	2024-04-08 21:37:54 -04:00
Wing Lian	2fa65b9599	ignore issues with calculating # params when printing (#1493 )	2024-04-08 11:04:22 -04:00
xzuyn	9430b6e868	Remove `validate_quantized_dora` (#1485 ) DoRA with quantized layers is supported with PEFT 0.10.0	2024-04-08 01:25:23 -04:00
Wing Lian	934fc851da	drop empty token from beginning if tokenizer has no bos_token (in the case of qwen) (#1490 )	2024-04-06 19:55:19 -07:00
NanoCode012	bda48f0150	fix: reduce sample_packing warning (#1484 )	2024-04-06 21:04:07 +09:00
NanoCode012	bf4cd67252	feat: validate sample packing requires flash_attention (#1465 ) * feat: validate sample packing requires flash_attention * fix: check for sdp_attn per suggestion * feat: add FA to tests	2024-04-05 12:47:32 +09:00
Wing Lian	05b0b7e8ca	add support for cohere chat template (#1478 )	2024-04-04 18:20:50 -07:00
Wing Lian	87ca3f98c6	don't use deepspeed or fsdp when merging loras (#1479 )	2024-04-04 18:20:32 -07:00
Wing Lian	e0fcef403f	refactor utils.data module for line count linter (#1476 )	2024-04-04 16:33:42 -07:00
NanoCode012	c2b64e4dcf	Feat: update doc (#1475 ) [skip ci] * feat: update doc contents * chore: move batch vs ga docs * feat: update lambdalabs instructions * fix: refactor dev instructions	2024-04-04 13:43:40 +09:00
Hamel Husain	5760099bd4	fix toc	2024-04-03 12:05:49 -07:00
Wing Lian	5aa50974ce	Pretrain multipack v2 (#1470 )	2024-04-02 05:42:16 -07:00
James Melvin Ebenezer	cae608f587	Added pip install ninja to accelerate installation of flash-attn (#1461 ) * Added pip install ninja to accelerate installation of flash-attn * doc: cleanup	2024-04-02 17:36:41 +09:00
Nick Doiron	586bd8d221	fix pretraining_ on odd datasets (#1463 ) * can configure name of split of pretraining dataset * streaming data and dataset map * text column customized * allow text_column to be set in pretrain * pretrain type * load a bit of the dataset * fix dataset where splits have separate configs * ok name param here is the config * whitespace	2024-04-01 20:48:59 -07:00
Hamel Husain	86b7d22f35	Reorganize Docs (#1468 )	2024-04-01 08:00:52 -07:00
Wing Lian	0b103775ad	reduce verbosity of the special tokens (#1472 )	2024-04-01 21:47:27 +09:00
NanoCode012	946b497c3f	feat: add deepspeed 3 with cpuoffload (#1466 ) * feat: add deepspeed 3 with cpuoffload * make bf16 explicit, add param only offload variant --------- Co-authored-by: Wing Lian <wing.lian@gmail.com>	2024-04-01 21:42:52 +09:00
Wing Lian	0ddfb24fcf	LISA (#1469 ) * add lisa support * fix default and fix attribute traversal for layers * improve lisa callback logging * fix LISA by ensuring params are not frozen during __init__ * example config for lisa --------- Co-authored-by: Aman Karmani <aman@tmm1.net>	2024-04-01 04:54:53 -07:00
Wing Lian	89134f2143	make sure to install causal_conv1d in docker (#1459 )	2024-03-29 16:43:25 -04:00
Wing Lian	6086be85f7	qwen2_moe support w multipack (#1455 )	2024-03-29 11:04:53 -04:00

1 2 3 4 5 ...

1421 Commits