axolotl

Author	SHA1	Message	Date
Salman Mohammadi	bdb0f97082	adding 'reward_processing_classes'	2025-02-05 18:18:42 +00:00
Salman Mohammadi	65b6519447	adding 'reward_processing_classes'	2025-02-05 18:13:05 +00:00
Wing Lian	a1958b09de	seperately include max_completion_len	2025-02-05 13:01:52 -05:00
Salman Mohammadi	b8f258817e	adding reward fn verification	2025-02-05 13:30:02 +00:00
Wing Lian	753146b458	max_length moved to reward config	2025-02-04 11:06:26 -05:00
Wing Lian	d683c50113	fix config cls	2025-02-04 11:06:26 -05:00
Wing Lian	234cd8311e	fix failure case in prompter loading	2025-02-04 11:06:26 -05:00
Wing Lian	f9893e3842	fix dpo config and add use_logits_to_keep	2025-02-04 11:06:26 -05:00
Wing Lian	ac1ebc58a8	add support for num_generations	2025-02-04 11:06:25 -05:00
Wing Lian	56f3b9f20f	bump pydantic to support vllm	2025-02-04 11:06:25 -05:00
Wing Lian	2c1376d8c4	don't shrink embeddings unless told to	2025-02-04 11:06:25 -05:00
Wing Lian	3c7517fd55	add support for passing map kwargs to dataset map in rl	2025-02-04 11:06:25 -05:00
Wing Lian	1e94d7ef65	more fixes to get grpo working	2025-02-04 11:06:25 -05:00
Wing Lian	cfc7fe0df2	remove ununsable args kwargs	2025-02-04 11:06:25 -05:00
Wing Lian	3c4fe478cf	be nice with self.cfg.dataset_processes	2025-02-04 11:06:25 -05:00
Wing Lian	c810599c66	order matters	2025-02-04 11:06:24 -05:00
Wing Lian	300ffc2cb6	make it a dataclass	2025-02-04 11:06:24 -05:00
Wing Lian	b1c4711145	load the class from strat	2025-02-04 11:06:24 -05:00
Wing Lian	d155849e2c	use correct builder	2025-02-04 11:06:24 -05:00
Wing Lian	626db6cb84	collator for grpo and prompt loader	2025-02-04 11:06:24 -05:00
Wing Lian	79159b4871	support custom module prompt strategy for rl	2025-02-04 11:06:24 -05:00
Wing Lian	704ddd6ff1	honor skip prepare for rl	2025-02-04 11:06:24 -05:00
Wing Lian	54b0d3d0e8	passthrough dataset parser for dpo/grpo	2025-02-04 11:06:23 -05:00
Wing Lian	59ad21f2de	refactor a bit for better grpo support	2025-02-04 11:06:23 -05:00
Wing Lian	57264b6491	respect dotenv for cli	2025-02-04 11:06:23 -05:00
Wing Lian	d495e41ba1	refactor dpo trainer into own module	2025-02-04 11:06:23 -05:00
Wing Lian	6067fe6c28	upgrade trl to 0.14.0	2025-02-04 11:06:23 -05:00
NanoCode012	a620d481e2	fix: drop long seq even if not sample packing (#2211 ) * fix: drop long seq even if not sample packing * fix: logging import * fix: cfg passed being none * fix: try to fix logging * fix: refactor call to not use accelerate log * fix: try to fix circular import issue * fix: don't drop when skip prepare * chore: remove duplicate line * fix: update warning to mention that sequences will be trimmed * fix: do not drop seq if input_ids don't exist * fix: increase RM unittest sequence length to reduce trim warnings * fix: solve conflicts * fix: default min_seq_len in case of None	2025-02-04 09:43:35 -05:00
Wing Lian	158330ab60	[feature] sweeps (#2171 )	2025-02-01 21:11:18 -05:00
Wing Lian	80e1468b8d	better handling of multipack dataset length (#2296 )	2025-02-01 21:10:34 -05:00
Wing Lian	a20f17689b	set MODAL_IMAGE_BUILDER_VERSION=2024.10 to 2024.10 to test latest builder (#2302 ) * set MODAL_IMAGE_BUILDER_VERSION=2024.10 to 2024.10 to test latest builder * chore: lint * remove fastapi and pydantic extras	2025-01-31 20:19:20 -05:00
Wing Lian	78ce268848	KD Trainer w logprobs (#2303 ) * refactor trainer to prevent circular dependencies later fix loader default KD dataset loading and KD with logprobs filter bad rows make batch smaller handle padding/collation for KD datasets make it work flipped the slice cross entropy loss coefficient during KD make sure to multiply against the correct loss chore: lint triton wip no where support v2 trial no torch.exp inside triton kernel no log etc no torch.tensor v3 fix kwarg don't use triton for now better rescaling for temperatures hash for temperature too use kd_alpha in the correct loss method fix kd loss so it's causal (fixes repeating tokens) var naming and add todo chore: lint refactor so we can easily add new loss functions add license block remove references to triton kd for now handle token/logprob shifting support for custom trainer classes from plugins refactor kd chat template loader move more things to kd plugin remove moved class from import make plugin setup concise increase logging around loading plugins add copyrights remove duplicate code more info on preprocess for kd and fix import be a bit pickier about loading dynamic prompt strategies kd sample packing make loss torch script compat support streaming for processing sft datasts? improve iterable support ensure that batch vs single is done properly tweak check for batched prompt data reward can use same batch check fix reward trainer calls for tokenization improve check for batched reward model doesn't work well with batched add kd trainer e2e test linting rename test files so it gets picked up make the kd e2e fit in vram for ci and add lora version set lora_dropout explicitly lower lr make sure to set tokenizer from l3 70b and save safetensors make sure to use the correct tokenizer fix adapter model check make sure to use tensorboard to capture loss for checks chore: lint chore: lint improve logprob masking and shift in trainer more fixes try tests for kd on l40s don't shift student logits for kd no batching for kd chat templates make sure to truncate logprobs if there are more than top_k change up logic so we always truncate to top_k use iter instead of tuple fix finding the top-k rather than assuming first position has the correct val apply z-score scaling to kd kd loss needs to be calculated in full precision Always re-normalize teacher distribution various fixes * support for configurable top-k/softmax ordering * add attribute check for filter rows and lint * fix logic * handle none case for conversion to int * fix student logit off by one * set kd_temp to 1.0 for test loss * address PR feedback	2025-01-31 20:18:52 -05:00
NanoCode012	d425d5d3c3	fix: add warning for invalid eval_steps or save_steps (#2298 )	2025-01-31 08:58:25 -05:00
Wing Lian	cf17649ef3	Misc fixes 20250130 (#2301 ) * misc fixes for garbage collection and L40S w NCCL P2P * patch bnb fix for triton check * chore: lint * change up import * try patching differently * remove patch for bnb fix for now * more verbose checks and tweak train loss threshold	2025-01-31 08:58:04 -05:00
Dan Saunders	6f294c3d8d	refactor README; hardcode links to quarto docs; add additional quarto doc pages (#2295 ) * refactor README; hardcode links to quarto docs; add additional quarto doc pages * updates * review comments * update --------- Co-authored-by: Dan Saunders <dan@axolotl.ai>	2025-01-30 12:49:21 -05:00
Wing Lian	6f713226dd	make save_safetensors: true the default (#2292 ) * make save_safetensors: true the default * revert change to model output check	2025-01-30 11:48:48 -05:00
Wing Lian	1063d82b51	match the cuda version for 2.4.1 build w/o tmux (#2299 )	2025-01-30 11:46:09 -05:00
salman	ac471a697a	updating to fused (#2293 )	2025-01-30 11:45:56 -05:00
Wing Lian	8779997ba5	native support for modal cloud from CLI (#2237 ) * native support for modal cloud from CLI * do lm_eval in cloud too * Fix the sub call to lm-eval * lm_eval option to not post eval, and append not extend * cache bust when using branch, grab sha of latest image tag, update lm-eval dep * allow minimal yaml for lm eval * include modal in requirements * update link in README to include utm * pr feedback * use chat template * revision support * apply chat template as arg * add wandb name support, allow explicit a100-40gb * cloud is optional * handle accidental setting of tasks with a single task str * document the modal cloud yaml for clarity [skip ci] * cli docs * support spawn vs remote for lm-eval * Add support for additional docker commands in modal image build * cloud config shouldn't be a dir * Update README.md Co-authored-by: Charles Frye <cfrye59@gmail.com> * fix annotation args --------- Co-authored-by: Charles Frye <cfrye59@gmail.com>	2025-01-30 11:34:02 -05:00
Eric Tang	268543a3be	Ray Train Axolotl Integration (#2251 ) * current not clean working version move torch trainer to do_cli update code with config changes and clean up edit config cleanup add run name to trainer * address comments * use axolotl train in multigpu tests and add ray tests for multi-gpu * accelerate uses underscores for main_process_port arg * chore: lint * fix order of accelerate args * include ray train in docker images * current not clean working version move torch trainer to do_cli update code with config changes and clean up edit config cleanup add run name to trainer * address comments * use axolotl train in multigpu tests and add ray tests for multi-gpu * accelerate uses underscores for main_process_port arg * chore: lint * fix order of accelerate args * include ray train in docker images * fix bf16 resolution behavior * move dtype logic * x Signed-off-by: SumanthRH <sumanthrh@anyscale.com> * rename Signed-off-by: SumanthRH <sumanthrh@anyscale.com> * add to sidebar Signed-off-by: SumanthRH <sumanthrh@anyscale.com> * Apply suggestions from code review Co-authored-by: Eric Tang <46737979+erictang000@users.noreply.github.com> * Update docs/ray-integration.qmd Co-authored-by: Eric Tang <46737979+erictang000@users.noreply.github.com> * pre-commit fixes Signed-off-by: SumanthRH <sumanthrh@anyscale.com> * use output_dir instead of hardcoded saves path Co-authored-by: NanoCode012 <kevinvong@rocketmail.com> * bugfix storage dir * change type\ for resources_per_worker --------- Signed-off-by: SumanthRH <sumanthrh@anyscale.com> Co-authored-by: Wing Lian <wing@axolotl.ai> Co-authored-by: SumanthRH <sumanthrh@anyscale.com> Co-authored-by: Sumanth R Hegde <39546518+SumanthRH@users.noreply.github.com> Co-authored-by: Wing Lian <wing.lian@gmail.com> Co-authored-by: NanoCode012 <kevinvong@rocketmail.com>	2025-01-29 00:10:19 -05:00
salman	54dd7abfc1	Process reward models (#2241 ) * adding model_cfg to set num_labels * using a num_labels field instead * linting * WIP stepwise prompt tokenizer * this should work? * trainer working? * pushing to runpod * fixing saving * updating conf * updating config, adding docs * adding stepwise supervision docpage * updating tests * adding test for dataset * fixing tests * linting * addressing some comments * adding additional cfg fields support * updating tests, fixing cfg * fixing tests * updating loss * Update test_process_reward_model_smollm2.py * updating loss values and seed * dumb pre-commit	2025-01-29 00:08:33 -05:00
salman	c071a530f7	removing 2.3.1 (#2294 )	2025-01-28 23:23:44 -05:00
mashdragon	c015a76a23	Num epochs float (#2282 ) [skip ci] * Change num_epochs type to float * Handle float value for num_epochs in trainer.py	2025-01-28 23:23:26 -05:00
NanoCode012	067b442596	chore: refactor SaveModelCallback to stop handle fractional save_steps (#2291 ) [skip ci]	2025-01-28 23:22:10 -05:00
Wing Lian	0b52f06227	bump bnb to 0.45.1 (#2289 ) [skip ci]	2025-01-28 23:21:25 -05:00
Wing Lian	887513285d	support for custom lr groups for non-embedding modules (#2213 ) * support for custom lr groups for non-embedding modules invert name check for group modules include lr_groups in training args additional conditional for creating optimizer fix regular params as w weight decay fix lookup and add docs * address pr feedback	2025-01-24 12:56:28 -05:00
Wing Lian	20620771f1	Pretrain multipack (#2278 ) * fix for pretrain with packing * fix model name and loss expected * make sure to check with micro batch size for pretraining * change loss threshholds based on parametrization * make tests smaller for CI * fix pretrain packing * fix pretrain packing test * address pr feedback	2025-01-24 12:55:20 -05:00
NanoCode012	6086162488	chore(doc): improve explanation for _steps and _strategy (#2270 )	2025-01-24 10:07:02 -05:00
mashdragon	b2774af66c	Take `split` param from config in all load_dataset instances (#2281 )	2025-01-24 10:06:50 -05:00
NanoCode012	74f9782fc3	chore(doc): fix explanation on gcs creds retrieval (#2272 )	2025-01-24 10:05:58 -05:00

1 2 3 4 5 ...

1867 Commits