candle

mirror of https://github.com/huggingface/candle.git synced 2025-06-16 18:48:51 +00:00

Author	SHA1	Message	Date
ivarflakstad	0c09d10f32	Improve metal buffer usage (#1807 ) * Improve metal buffer usage * Clone cpu storage when loading to reduce wait_until_complete calls * Use powers of two for buffer sizes so reuse is more likely. * Select best available buffer by size. * Add count to MetalStorage -> can use buffer with different size Co-authored-by: Chris Fleetwood <christopher.fleetwood@huggingface.co> * Simplify new buffer creation without blit copy. Revert &[] -> Vec * Add documentation on newBufferWithBytes safety / synchronization * Drop unused buffers after command buffer is done syncing. --------- Co-authored-by: Chris Fleetwood <christopher.fleetwood@huggingface.co>	2024-03-07 09:42:34 +01:00
Laurent Mazare	8a99cf7dd2	Add a flag to select the dtype used in metavoice. (#1805 )	2024-03-05 12:16:00 +01:00
Laurent Mazare	bd9ab9bc04	Add a cuda kernel for dequantizing q8_0. (#1804 )	2024-03-05 09:50:37 +01:00
Laurent Mazare	8cc0a183ba	Speaker embeddings computation for metavoice. (#1800 ) * Speaker embeddings computation for metavoice. * Compute the speaker embeddings.	2024-03-04 14:13:01 +01:00
Laurent Mazare	6530932285	Add the new models to the main readme. (#1797 )	2024-03-03 16:25:14 +01:00
Jiayu Liu	924ccae30c	Add an initial Segformer implementation (#1617 ) * add segformer * Make the id2label field optional. --------- Co-authored-by: laurent <laurent.mazare@gmail.com>	2024-03-03 16:01:46 +01:00
Laurent Mazare	60dc72b96b	More metavoice tweaks. (#1796 )	2024-03-03 15:05:25 +01:00
Laurent Mazare	20abb72fec	Normalize loudness of the generated audio (#1795 ) * Normalize loudness of the generated audio. * Lints. * One more lint. * Avoid running the bs1770 tests. * Another attempt at discarding doc comments. * Also normalize the loudness in the encodec example.	2024-03-03 14:00:42 +01:00
Laurent Mazare	ca5d727ba2	Use the same padding in metavoice as in the python version. (#1794 )	2024-03-03 12:04:48 +01:00
Laurent Mazare	09e0148cce	Tweaks to run metavoice on metal (#1792 ) * Enable tanh + tweak conv-transpose. * Run the encodec decoding on cpu. * Clippy fixes.	2024-03-03 07:46:44 +01:00
Laurent Mazare	de11623752	Metavoice position fix (#1791 ) * Add the metavoice transformer. * Sketch the speaker-encoder module. * Adding to the metavoice model. * Start adding the metavoice example. * Get some logits out. * Load the second stage model. * Get the second step to run. * Tweak the example. * Add encodec tilting. * Glue the different bits together. * Fix a shape issue. * Use a constant. * BPE tokenization. * Fix the position index in metavoice.	2024-03-02 21:00:35 +01:00
Laurent Mazare	21f1d04976	Add the instruction finetuned gemma variants. (#1790 )	2024-03-02 18:56:59 +01:00
Laurent Mazare	4fff5b51f5	Metavoice - first cut (#1717 ) * Add the metavoice transformer. * Sketch the speaker-encoder module. * Adding to the metavoice model. * Start adding the metavoice example. * Get some logits out. * Load the second stage model. * Get the second step to run. * Tweak the example. * Add encodec tilting. * Glue the different bits together. * Fix a shape issue. * Use a constant. * BPE tokenization. * Add a warning.	2024-03-02 18:50:01 +01:00
Laurent Mazare	314630638d	Rustfmt fix. (#1788 )	2024-03-02 10:35:07 +01:00
Frkri	3e3def4134	Update StableLM config (#1787 )	2024-03-02 09:56:57 +01:00
Jack Shih	6980774a91	fix rwkv example eos token (#1785 )	2024-03-01 10:22:28 +01:00
Laurent Mazare	64d4038e4f	Mention rwkv v6 in the readmes. (#1784 )	2024-03-01 08:58:30 +01:00
Jani Monoses	979deaca07	EfficientVit (MSRA) model (#1783 ) * Add EfficientVit (Microsoft Research Asia) model. * Mention models in README	2024-03-01 08:53:52 +01:00
Jack Shih	b485e4b6ee	add models of rwkv v6 and quantized rwkv v6 (#1781 ) * add models of rwkv v6 and quantized rwkv v6 * fix ci clippy fail	2024-03-01 08:37:56 +01:00
laurent	2c95b7394a	Handle Q5_0 and Q5_1 quants in cuda.	2024-02-29 10:54:01 +01:00
Laurent Mazare	4fd00b8900	Add the StarCoder2 model. (#1779 ) * Add the StarCoder2 model. * Add the example code and get things to work. * And also tweak the readme.	2024-02-28 21:02:41 +01:00
Laurent Mazare	57267cd536	Add a flag to force running the quantized model on CPUs. (#1778 ) * Add a flag to force running the quantized model on CPUs. * Add encodec to the readme.	2024-02-28 14:58:42 +01:00
Laurent Mazare	60ee5cfd4d	Support more modes in the encodec example. (#1777 ) * Support more modes in the encodec example. * Remove the old encodec model from the musicgen bits.	2024-02-28 09:22:33 +01:00
Laurent Mazare	56e44aabe3	Make some dependencies optional in the examples. (#1776 )	2024-02-28 07:17:03 +01:00
Laurent Mazare	d0aca6c3c6	Encodec encoding demo. (#1775 )	2024-02-28 06:49:03 +01:00
Laurent Mazare	15e8644149	Apply dilations in the encodec model. (#1772 ) * Apply dilations in the encodec model. * Add some encoding bits.	2024-02-27 23:26:35 +01:00
Laurent Mazare	0c49e95dfb	Encodec model. (#1771 ) * Encodec model. * Fixes. * Add the padding functions. * Get the LSTM bit to work. * Get the encodec model to generate some tokens (decoder only for now). * Minor tweak. * Minor tweak.	2024-02-27 22:59:40 +01:00
Laurent Mazare	205767f9de	Avoid tensor copying in the quantized example. (#1770 )	2024-02-27 20:32:30 +01:00
Laurent Mazare	5e526abc8c	Bump the version number to 0.4.1. (#1768 ) * Fix the block size for some cuda kernels. * Bump the version number to 0.4.1.	2024-02-27 14:19:59 +01:00
Laurent Mazare	6400e1b0a0	Fix the block size for some cuda kernels. (#1767 )	2024-02-27 14:08:33 +01:00
Laurent Mazare	32544a2ad6	Add an option to split the prompt. (#1766 )	2024-02-27 11:24:11 +01:00
Laurent Mazare	badf886583	Cuda kernel for dequantizing q8k. (#1760 ) * Cuda kernel for dequantizing q8k. * Clippy lints.	2024-02-26 08:42:44 +01:00
Jack Shih	918136ba46	add quantized rwkv v5 model (#1743 ) * and quantized rwkv v5 model * Integrate the quantized rwkv model in the initial example. --------- Co-authored-by: laurent <laurent.mazare@gmail.com>	2024-02-25 21:43:40 +01:00
Laurent Mazare	1a6043af51	Tweak the VarMap set type. (#1758 )	2024-02-25 20:50:08 +01:00
Laurent Mazare	2f22afd80e	Cuda acceleration for quantized model. (#1754 ) * Boilerplate for the quantized cuda support. * More basic cuda support. * More cuda quantization (quantize on cpu for now). * Add the dequantization bit. * Start adding some dedicated cuda kernels from llama.cpp. * Move the kernel code. * Start interfacing with the kernel. * Tweak the kernel launch params. * Bugfix for quantized metal. * Fix some clippy lints. * Tweak the launch parameters. * Tweak cuda basics to perform a quantized matmul. * Perform the dequantization on the cpu + use cublas for matmul. * Add the dequantization kernel. * Test the qmatmul. * More kernels. * Matmul-vec kernel. * Add a couple kernels. * More dequantization kernels.	2024-02-25 18:11:47 +01:00
Laurent Mazare	8d04f70f4d	Fix the eos token for gemma. (#1753 )	2024-02-24 11:07:02 +01:00
Laurent Mazare	eeb7e2b683	Apply rustfmt to the newly added tests. (#1749 )	2024-02-23 06:48:28 +01:00
Sacha Arbonel	11ea7aac4d	tests (#1724 )	2024-02-23 06:35:46 +01:00
Daniel Varga	32eb56d6b3	Fix typo in README (#1740 )	2024-02-22 12:35:26 +01:00
Laurent Mazare	28057781aa	Make the cache for the llama model explicit too. (#1745 )	2024-02-22 12:04:33 +01:00
laurent	544018b6d0	Explicit caching in llama2.c.	2024-02-22 10:22:03 +01:00
Laurent Mazare	c753f72c85	Support for attention bias in gemma + refactor things a bit. (#1744 ) * Support for attention bias in gemma + refactor things a bit. * Fix the cuda tests.	2024-02-22 09:35:28 +01:00
Kirpal Grewal	8013b50829	Add grads for interpolate1d (#1742 ) * add backprop for interpolate1d * fix clippy lint * correct fix clippy lint	2024-02-22 08:44:01 +01:00
Laurent Mazare	45d5322d62	Add the Gemma models. (#1741 ) * Add the Gemma models. * Add the gemma example. * Adapt the RmsNorm. * Get the 2b model to work. * 7b support. * Use the config head dim. * Yet another fix. * Make the matrixes contiguous. * Also get the 7b model to work. * And add to the readme.	2024-02-21 22:02:50 +01:00
Laurent Mazare	a2cb2edead	Add a couple backtraces on cpu errors. (#1738 )	2024-02-20 19:54:13 +01:00
Laurent Mazare	fc67d878bb	Bugfix for conv-transpose1d (#1734 ) * Add a currently broken test. * Bugfix + fix test.	2024-02-19 09:04:49 +01:00
Laurent Mazare	3ba37443e5	Bugfix for applying the bias in conv1d-transpose. (#1732 )	2024-02-18 22:51:20 +01:00
Laurent Mazare	1fb728772d	Support for groups in conv-transpose1d. (#1731 ) * Groups support in conv-transpose-1d. * Remove dangling file.	2024-02-18 21:28:07 +01:00
Laurent Mazare	cb86b0c82c	Fix float unpickling. (#1730 )	2024-02-18 19:33:55 +01:00
Laurent Mazare	6284ad784c	Module implementation for options. (#1728 )	2024-02-18 14:12:55 +01:00

1 2 3 4 5 ...

1830 Commits