candle

mirror of https://github.com/huggingface/candle.git synced 2025-06-15 18:28:24 +00:00

Author	SHA1	Message	Date
Laurent Mazare	be9c200cbb	Expose the t5 config fields + allow t5-large. (#1987 )	2024-04-01 20:58:34 +02:00
Laurent Mazare	8ad12a0e81	Add some examples using the MT5 variants. (#1963 )	2024-03-29 18:09:29 +01:00
Laurent Mazare	c5092f2c29	Add a couple t5 models. (#1958 )	2024-03-28 17:58:06 +01:00
Laurent Mazare	37c539f2b7	Helper function to load sharded safetensors files (#1481 ) * Fix the quantized mistral example. * Add a helper function to load sharded safetensors weights. * Use the sharded loader.	2023-12-25 21:49:21 +01:00
Juarez Bochi	18d30005c5	Add support to UL2 model family (#1300 ) * Add support to UL2 model family * Update docs with UL2 * Create ActivationWithOptionalGating to avoid polluting activations * Also refactor quantized t5 * Remove useless conversion * Revert Activation::NewGelu name change * Remove useless return * Apply rustfmt and clippy recommendations * Reuse t5::ActivationWithOptionalGating in quantized version * (cosmetic change) use a match rather than ifs + avoid early returns. --------- Co-authored-by: Laurent <laurent.mazare@gmail.com>	2023-11-09 18:55:09 +01:00
Juarez Bochi	508f811b93	Add support for MADLAD400 (#1285 ) * Add support for madlad * Add support for quantized MADLAD	2023-11-07 05:35:37 +01:00
Laurent Mazare	29c7f2565d	Add some reinforcement learning example. (#1090 ) * Add some reinforcement learning example. * Python initialization. * Get the example to run. * Vectorized gym envs for the atari wrappers. * Get some simulation loop to run.	2023-10-14 16:46:43 +01:00
Laurent Mazare	890d069092	Self-contained safetensor wrappers (#946 ) * Self-contained safetensor wrappers. * Use the new safetensor container in varbuilders.	2023-09-23 20:39:52 +01:00
Laurent Mazare	aa8ec06fd2	Add the t5-xxl version. (#924 )	2023-09-21 14:48:13 +01:00
Laurent Mazare	9b24d89d2d	Tracing mode for T5. (#913 ) * Tracing mode for T5. * Tracing for the linear layer.	2023-09-20 15:03:35 +01:00
Juarez Bochi	8696f64bae	Fix T5 kv cache (#899 ) * Fix T5 kv cache * Add argument for decoder prompt * Fix range	2023-09-19 20:36:15 +01:00
Juarez Bochi	1542e92629	T5: Add option to override use_cache from config (#892 ) * Add option to override use_cache from config * Disable cache by default and cleanup code	2023-09-18 20:20:21 +01:00
Laurent Mazare	7f65af1f0d	Avoid re-encoding the input in the T5 example. (#875 )	2023-09-17 10:25:54 +01:00
Laurent Mazare	eeb54716dd	Tweaks for the T5 example. (#874 )	2023-09-17 10:05:15 +01:00
Laurent Mazare	1a276b5da7	Add a KV cache to T5. (#873 ) * Add a KV cache to T5. * Suggest using release mode. * Use the kv cache in decoding. * Add a comment.	2023-09-17 08:00:45 +01:00
Juarez Bochi	3e49f8fce5	Implement T5 decoding (#864 ) * Load t5 decoder * Run enc, dec, and lm head, but no cross attn * Cross-attention over key_value_states * New arg for decoder input ids * Add mask, don't forward position biases through decoder * Update t5 examples * Clippy + rustfmt	2023-09-15 22:05:12 +02:00
Laurent Mazare	31ab2ddaeb	Remove the padding. (#838 )	2023-09-13 13:00:59 +01:00
Laurent Mazare	3e94324012	Add some sentence similarity part to the t5 example. (#835 ) * Add some sentence similarity part to the t5 example. * Clippy fix.	2023-09-13 10:44:02 +01:00
Laurent Mazare	e4553fb355	T5 tweaks (#831 ) * Use default values rather than options. * Avoid exposing the device field. * More tweaks.	2023-09-13 07:37:04 +01:00
Juarez Bochi	9daa6dbe87	Extract T5 module and add main function to use it (#829 ) * Extract t5 out of musicgen * Add main for t5 module	2023-09-13 07:14:05 +01:00

20 Commits