axolotl

Files

salman 54dd7abfc1 Process reward models (#2241 )

* adding model_cfg to set num_labels

* using a num_labels field instead

* linting

* WIP stepwise prompt tokenizer

* this should work?

* trainer working?

* pushing to runpod

* fixing saving

* updating conf

* updating config, adding docs

* adding stepwise supervision docpage

* updating tests

* adding test for dataset

* fixing tests

* linting

* addressing some comments

* adding additional cfg fields support

* updating tests, fixing cfg

* fixing tests

* updating loss

* Update test_process_reward_model_smollm2.py

* updating loss values and seed

* dumb pre-commit

2025-01-29 00:08:33 -05:00

messages

wip add new proposed message structure (#1904 )

2024-10-13 12:15:18 -04:00

__init__.py

move shared pytest conftest to top level tests (#2099 ) [skip ci]

2024-11-22 15:05:42 -05:00

conftest.py

fix: use apply_chat_template to find turn boundaries and allow tool_calling field (#2179 ) [skip ci]

2024-12-17 16:42:21 -05:00

test_alpaca.py

fix broken linting (#1541 )