various tests fixes for flakey tests (#2110)

* add mhenrichsen/alpaca_2k_test with revision dataset download fixture for flaky tests

* log slowest tests

* pin pynvml==11.5.3

* fix load local hub path

* optimize for speed w smaller models and val_set_size

* replace pynvml

* make the resume from checkpoint e2e faster

* make tests smaller
This commit is contained in:
Wing Lian
2024-12-02 17:28:58 -05:00
committed by bursteratom
parent b0fbd4d11d
commit c0c53eb62f
13 changed files with 78 additions and 44 deletions

View File

@@ -52,6 +52,7 @@ class TestReLoraLlama(unittest.TestCase):
],
"warmup_steps": 15,
"num_epochs": 2,
"max_steps": 51, # at least 2x relora_steps
"micro_batch_size": 4,
"gradient_accumulation_steps": 1,
"output_dir": temp_dir,