axolotl

Author	SHA1	Message	Date
Wing Lian	ee262818ef	fix evals (#447 )	2023-08-20 23:39:42 -04:00
Wing Lian	9d629d8bff	gracefully handle empty input (#442 )	2023-08-20 09:18:18 -04:00
Wing Lian	d2e7f27240	support user defined prompters, pretokenized datasets in config, local parquet, local arrow files (#348 ) * support user defined prompters, pretokenized datasets in config, local parquet, local arrow files * fix user defined dataset types * fix for system prompts * fix tests * fix checks for parquet and arrow * aha moment that d.data_files isn't used * add documentation for ds_type to add support for parquet and arrow	2023-08-20 09:17:49 -04:00
Wing Lian	f733d0f31e	disable eval using multipack for now (#437 )	2023-08-19 10:35:04 -04:00
Wing Lian	008505c8ae	fix comma, not a tuple (#436 )	2023-08-19 00:57:40 -04:00
Wing Lian	b3f5e00ff5	use save_strategy from config if available (#434 ) * use save_strategy from config if available * update docs for save_strategy	2023-08-18 20:28:23 -04:00
Wing Lian	5247c5004e	set env for FSDP offload params (#433 )	2023-08-18 20:28:09 -04:00
Aman Gupta Karmani	06edf175ac	standardize attn hijack patches (#381 ) * split sdp attn into its own patch * sync xformers patch to follow shared format and be diffable * update flash-attn patch for 70B/GQA and inference using helper from flash-attn tests * speed up flash-attn inference * fix patch to check position ids and don't use multipack for evals * copy LlamaModel.forward and LlamaDecoderLayer.forward into monkeypatch * update forwards so we only calculate cu_seqlens once * enable eval dataloader using multipack again * fix the patch to work properly and work with FSDP --------- Co-authored-by: Wing Lian <wing.lian@gmail.com>	2023-08-18 12:54:16 -04:00
mhenrichsen	0a228479b3	adds color (#425 ) * adds color * chore: lint * fix for colorama --------- Co-authored-by: Wing Lian <wing.lian@gmail.com>	2023-08-18 10:59:43 -04:00
Wing Lian	1b7e8604bb	fix orca prompts (#422 )	2023-08-16 11:21:03 -04:00
NanoCode012	c01015f33f	Fix(config): Update handling of deepspeed config (#404 ) * Fix(config): Update handling of deepspeed config * feat: auto set deepspeed env if deepspeed passed * fix: update new deepspeed instructions	2023-08-16 01:22:43 +09:00
Wing Lian	da10af03e9	fix eval steps and strategy (#403 )	2023-08-15 07:28:50 -04:00
Wing Lian	85cf4f8e2c	better handling of empty input ids when tokenizing (#395 ) * better handling of empty input ids when tokenizing * Add warning if tokenizer resulted in empty result * fix len comparison for linter	2023-08-15 01:09:59 -04:00
Aman Karmani	2e22404d2d	add utils.data.prepare_dataset	2023-08-14 21:28:29 -07:00
Wing Lian	fc2d6be96d	use context manager to run things on rank0 before others (#397 )	2023-08-15 00:10:47 -04:00
Wing Lian	1687be6a35	don't use mask expansion for inference (#392 )	2023-08-14 20:52:54 -04:00
Gabriel Puliatti	3c2ad00d07	Feat(config): add max steps (#387 )	2023-08-14 11:19:29 -04:00
florian peyron	5d48a10548	Added "epoch" evaluation_strategy (#388 )	2023-08-14 10:59:23 -04:00
NanoCode012	73a0b6ead5	Feat(config): Add hub_strategy (#386 )	2023-08-14 07:12:55 -04:00
florian peyron	63fdb5a7fb	Error msg for sharegpt if conv has less than 2 msg (#379 )	2023-08-14 17:40:40 +09:00
Wing Lian	919246fbc1	don't pass rope_scaling kwarg if it's None (#383 )	2023-08-13 18:57:38 -04:00
Charles Goddard	15f6e57eaa	Fix crash when running without CUDA	2023-08-13 13:36:40 -07:00
Aman Gupta Karmani	094fc2c6e6	try to detect accelerate and only use device_map=None in that case (#373 )	2023-08-13 00:32:07 -04:00
Wing Lian	343ac84e5a	fix check for flash attn branching (#377 )	2023-08-12 22:48:08 -04:00
Aman Karmani	0c967279ce	remove unnecessary local variable	2023-08-13 01:58:39 +00:00
Aman Karmani	efb3b2c95e	simplify `load_tokenizer`	2023-08-12 18:55:06 -07:00
Aman Karmani	7b55fe6419	improve GPU logging to break out pytorch cache and system mem	2023-08-12 18:52:57 -07:00
Aman Karmani	e029ab34ea	quiet noise from llama tokenizer by setting pad token earlier	2023-08-12 18:31:40 -07:00
Aman Karmani	8cec513447	extract module for working with cfg	2023-08-12 18:25:27 -07:00
Aman Karmani	a13e45d548	fix DefaultDict.__or__	2023-08-13 01:15:50 +00:00
Wing Lian	2bb0b78975	Attention mask and position id fixes for packing (#285 ) * fix attetion mask with packing * set position ids and use block diagonal attn mask * fix expand mask for multiple batch items, make sure we pad position_ids * don't move masks to cpu * use multi pack dataloader w random sampler * add position_ids back * more fixes for dataloader integration * est total tokens, fix field loop * more fixes, position_ids seems broken * more fixes for sample packing * use distributed sampler, avoid accelerate prepare * use accelerator prepare for dataloader * fix for position_ids w packing * Update src/axolotl/utils/dataloader.py * validation for sample packing and doc * more fixes for 4k and optimizations * optimized expand mask fn * better handling of variance in multipack dataloader length and trainer hanging when it runs out of data * fix rounding of len of batches to int * better handling so that all devices have the same dataloader len * fix step calc for packing * pass sample packing efficiency to training args * add a test for the mask expansion for sequence packing * only process eval dataset for packing if not None * don't split batches when packing * weighted CE losses * weighted CEL fixes * limit packing to sequences of max seq len * seq_len_multiple for packing * make sure the chunk size is an int * sample_packing_seq_len_multiplier config * use cumulative seq len with var len flash attn v2 w packing * properly calculate max len * fix flash-attn, xformers, packing, support chatml * fix chatml system prompt for openorca, legacy tokenizer opts * add chatml * add unit tests for cum seq lens, add ability to build cu_seq_lens from positional ids, fix prompt test * fix test and pylint checks * more packing and dataset optimizations and fixes * filter w multiple cpus * more fixes and optimizations * fixes and go back to distributed sampler since batch sampler won't work * fix counts by accounting for num devices * fix steps calculation * previous accelerate is still most performant * add numba to requirements. * use custom distributed checks * fix sampler to prevent overfit w new epochs * let's not cleanup the cached datasets * calculate cum seq lens with pos_ids instead of mask, simplify packing params, fix distributed barrier * speed optimizations and set accelerate fsdp env vars * optimize dataset concatenation? * more optimizations for dataset handling * fix import for annotation * manual pre-commit fixes * another sum optimization and bug fix for calc steps * fix packing estimations * fix formatting * pylint problems * add back flash attention branch for handling unpacked sequences seperately * Address PR feedback * add optional sample packing config params to readme	2023-08-12 15:14:56 -04:00
Morgan McGuire	7019509daa	Add wandb_entity to wandb options, update example configs, update README (#361 ) * Update wandb_entity and add wandb descriptions * add wandb to config section * remove trailing whitespace for pre-commit hook * remove trailing whitespace for pre-commit hook --------- Co-authored-by: Morgan McGuire <morganmcguire@Morgans-MacBook-Pro.local> Co-authored-by: Wing Lian <wing.lian@gmail.com>	2023-08-12 12:17:11 -04:00
NanoCode012	96bd6ae1c4	Fix(model loading): Warn when model revision is passed to gptq (#364 ) * fix(model loading): warn when model revision is passed to gptq * chore: improve message	2023-08-13 01:16:59 +09:00
NanoCode012	e37d9358e6	Fix(message): Improve error message for bad format (#365 )	2023-08-13 01:16:18 +09:00
NanoCode012	b5212068ac	Feat: Add rope scaling (#343 ) * Feat: Add rope scaling * fix: move rope config	2023-08-13 00:50:15 +09:00
Aman Gupta Karmani	11ddccb80f	Merge pull request #356 from tmm1/load_model-args simplify `load_model` signature	2023-08-09 18:24:34 -07:00
Aman Karmani	718102271f	simplify load_model signature	2023-08-09 22:36:02 +00:00
Aman Karmani	e303d64728	log GPU memory usage	2023-08-09 18:26:28 +00:00
Wing Lian	176b888a63	ensure enable_input_require_grads is called on model before getting the peft model (#345 )	2023-08-06 18:13:10 -04:00
Jan Philipp Harries	3392270544	experimental llama 2 chat support (#296 ) * experimental llama 2 chat support * few small fixes * llama2_chat * small fix to follow original implementation * small fixes and added fixtures/tests * fix -mixed up inference and finetuning conversations * args - small fix * small fix * small adjustment and warning * fix with pre-commit --------- Co-authored-by: Jan Philipp Harries <jpdus@users.noreply.github.com>	2023-08-06 17:40:52 -04:00
ssmi153	10405b9995	Update XFormers Attention Monkeypatch to handle Llama-2 70B (GQA) (#339 ) * Fix XFormers attention for Llama-2 70B (GQA) Updated XFormers MonkeyPatch to handle GQA as used in Llama-2 70B. All the updated code is taken directly from the Transformers library: `07360b6c9c (diff-06392bad3b9e97be9ade60d4ac46f73b6809388f4d507c2ba1384ab872711c51)` from their llama_modeling.py file. * Catch configs without pretraining_tp * Whitespace bug fix Command had accidentally been moved out of if-else block. * pre-commit formatting fixes Thanks to @winglian	2023-08-06 11:09:04 -04:00
Jan Philipp Harries	c93655c0a3	Added Orca Mini prompt strategy (#263 ) * added Orca Mini prompt strategy * maybe this fixed precommit errors? * pre-commits passing --------- Co-authored-by: Jan Philipp Harries <jpdus@users.noreply.github.com>	2023-08-06 03:16:41 +09:00
Wing Lian	fe285430bc	optimize the iteration when tokenizeing large datasets (#332 )	2023-08-04 12:12:05 -04:00
Aman Karmani	2eda9e02a9	fix typo	2023-08-03 21:04:12 +00:00
Aman Karmani	78b9efb7f4	scope flash-attn+qlora fix correctly, scope to llama, add comment	2023-08-03 19:19:39 +00:00
Aman Karmani	312a9fad07	move flash-attn monkey patch alongside the others	2023-08-03 17:20:49 +00:00
Aman Karmani	248bf90f89	ensure flash-attn fixes happen in both adapter/lora modes, and use torch_dtype	2023-08-02 20:15:03 +00:00
Wing Lian	77085ea24e	qlora w flash attention fixes (#333 )	2023-08-01 23:26:16 -04:00
Wing Lian	db2a3586f3	add peft install back since it doesn't get installed by setup.py (#331 )	2023-07-31 16:31:53 -04:00
Wing Lian	3d4984b9a5	update prompts for open orca to match the paper (#317 ) fix the test for the updated system tokenizer	2023-07-22 13:49:11 -04:00

1 2 3 4 5 ...

322 Commits