candle

mirror of https://github.com/huggingface/candle.git synced 2025-06-14 18:06:36 +00:00

Author	SHA1	Message	Date
Laurent Mazare	e27b4700ad	Indexing with max-value results in zero/no-op. (#2940 ) * Indexing with max-value results in zero/no-op. * Add some testing. * Also adapt the metal kernels. * Another test. * Fix.	2025-05-03 11:36:31 +02:00
Laurent Mazare	8a19bb7df2	Bump the candle version to 0.9.1. (#2935 )	2025-05-01 10:08:16 +02:00
Laurent Mazare	e3db30021f	Support for "unbatched" rope. (#2926 ) * Support for (un)-batched rope. * Use 3d rope in the rope/ropei/rope_thd functions. * Get the CPU versions to work. * Fix the cuda version. * Adapt the metal side. * Fix the metal tests.	2025-04-27 15:12:02 +02:00
Laurent Mazare	fbaf0b0e32	Bump the crate version to 0.9.0. (#2924 )	2025-04-26 11:01:21 +02:00
Laurent Mazare	3827685524	Add the scatter op. (#2921 ) * Add the scatter op. * Backprop support. * Cuda support.	2025-04-25 21:46:58 +02:00
Laurent Mazare	a4c56a958e	Add the const-set op. (#2910 ) * Add the const-set op. * Cuda implementation. * Bugfix. * Metal cleanup. * Add the metal kernels. * Add some testing. * Finish the metal implementation. * Bump the version.	2025-04-19 10:07:02 +02:00
Laurent Mazare	ce5f8dd129	Check the bounds in the cuda indexing kernels. (#2908 ) * Check the bounds in the cuda indexing kernels. * Another check.	2025-04-18 20:08:17 +02:00
Laurent Mazare	e4e7b0b2da	Use cudarc 0.16. (#2900 ) * Use cudarc 0.16. * Allow for disabling event tracking. * Tweaks. * Bump the ug version. * And bump the candle version too.	2025-04-15 21:40:18 +02:00
Laurent Mazare	1d1d6d4fe6	Bump the crate version. (#2895 )	2025-04-14 15:52:11 +02:00
Laurent Mazare	d9198deb37	Im2col cuda optimization. (#2885 )	2025-04-13 10:07:53 +02:00
Laurent Mazare	19fb6dac1f	Bump the crate version. (#2881 )	2025-04-11 22:28:21 +02:00
Laurent Mazare	d9904a3baf	Update to cudarc 0.14 (breaking change). (#2858 ) * Start updating to cudarc 0.14. * Adapt a couple more things. * And a couple more fixes. * More tweaks. * And a couple more fixes. * Bump the major version number. * Proper module system for the cuda kernels. * Proper ptx loading. * Launch the sort kernel. * Custom op. * Start using the builder pattern. * More builder. * More builder. * Get candle-core to compile. * Get the tests to pass. * Get candle-nn to work too. * Support for custom cuda functions. * cudnn fixes. * Get flash attn to run. * Switch the crate versions to be alpha. * Bump the ug dependency.	2025-04-03 09:12:19 +02:00
Laurent Mazare	468d1d525f	Bump the crate version to 0.8.4. (#2808 )	2025-03-15 07:42:24 +01:00
Laurent Mazare	fd7f7242a1	Bump the crate version to 0.8.3 (#2772 ) * update to cudarc to v0.13.5 to support cuda 12.8 * Bump the crate version. --------- Co-authored-by: Michael McCulloch <michael.james.mcculloch@fastmail.com>	2025-02-15 15:54:48 +01:00
Laurent Mazare	236c35e578	Bump the caret version to 0.8.2. (#2703 )	2025-01-07 15:50:16 +01:00
Laurent Mazare	67cab7d6b8	Bump the crate version to 0.8.1. (#2662 )	2024-12-07 17:03:53 +01:00
Laurent Mazare	1a0f9ccf16	Import the ggml_cuda_dp4a function. (#2628 )	2024-11-19 03:41:34 +01:00
Laurent Mazare	9453cc3095	Bump the crate version to 0.8.0. (#2612 )	2024-11-12 14:11:46 +01:00
Laurent Mazare	6454597943	Improved launch config for layer-norm/rms-norm. (#2591 ) * Improved launch config for layer-norm/rms-norm. * Add more testing for the fused layer/rms norm kernels.	2024-11-04 10:42:18 +01:00
Laurent Mazare	3a3c48b14b	Bump the crate version to 0.7.2. (#2517 )	2024-09-29 10:56:50 +02:00
Laurent Mazare	8097559c1a	Move the candle version to 0.7.1. (#2495 )	2024-09-22 20:44:39 +02:00
Laurent Mazare	c2fca0ca11	Bump the crate version. (#2491 )	2024-09-21 15:13:12 +02:00
Laurent Mazare	6070278a31	Bump the version to 0.6.1. (#2438 )	2024-08-22 09:23:52 +02:00
Laurent Mazare	f65e90e7ef	Bump the crate version. (#2248 )	2024-06-05 15:49:15 +02:00
Laurent Mazare	1df2bddccf	Add the layernorm specialized op. (#2212 ) * Add the layernorm cuda kernels. * Dedicated layer norm op. * Add the slower variant. * Plug the cuda implementation. * Add the metal variant. * Add a dedicated test. * Bugfix.	2024-05-24 15:58:01 +02:00
Laurent Mazare	6f0b807ffd	More efficient cuda implementation for ConvTranspose1d. (#2211 ) * More efficient cuda implementation for ConvTranspose1d. * Small tweak.	2024-05-24 11:05:43 +02:00
Laurent Mazare	89f53b9d7b	Bump the version number to 0.5.1. (#2155 ) * Bump the version number to 0.5.1. * Fix clippy lints for 1.78. * More clippy fixes.	2024-05-03 11:17:05 +02:00
MilkFather	3bbb88fcb4	Fix sigmoid gradient calculation and move sigmoid into a specialized op (#2114 ) * add sigmoid op * small fix * add as a method on `Tensor` * implement gradient calculation for sigmoid * add sigmoid tests * we should have a specialized op for this * fix clippy * fix clippy 2 * Revert all previous commits in favor of a `CustomOp` based solution * use `CustomOp1` implementation * fix rustfmt * experimental add metal impl * add cuda kernel impl * fix fmt * Add a test + reduce some cuda duplication. --------- Co-authored-by: laurent <laurent.mazare@gmail.com>	2024-04-29 11:04:43 +02:00
Laurent Mazare	eb26e2467e	Add the cuda dequantize f16 kernels. (#2137 ) * Add the cuda dequantize f16 kernels. * Expose the cuda kernels. * Add some testing + fix. * Test the other cases too. * A few more tests. * Add an environment variable to enable the dequantize f16 + matmul behavior.	2024-04-28 20:05:05 +02:00
Laurent Mazare	96a48e5cc4	Add argsort. (#2132 ) * Add the argsort cuda kernels. * CPU version of arg-sort. * Hook the cuda kernel + rework the cpu bits. * Add some dedicated test. * Working cuda kernel. * Metal kernel. * Metal adjustments. * Bugfix. * Use the fast rope in qwen. * Rework the expert selection in qwen.	2024-04-27 20:17:35 +02:00
Laurent Mazare	8de0ce6cba	Add more QMMV cuda kernels. (#2077 ) * Add more QMMV cuda kernels. * Enable the new kernels. * Adapt the testing.	2024-04-18 08:36:43 +02:00
Laurent Mazare	2817643db9	Add the mmv kernels for small batch sizes. (#2075 ) * Add the mmv kernels for smaller sizes. * Support more mmv kernels. * Use the new kernels. * Fix the call. * Silly fix. * Improve the testing. * Fix for dmmv. * Add another dedicated test for the batching mmv.	2024-04-16 21:30:51 +02:00
Laurent Mazare	f7d5bf5b97	Faster kernels for quantized matmul on cuda (#2060 ) * Hook the quantized matmul cuda kernels. * Add a (currently broken) test. * Kernel fixes. * Fix by transposing the rhs matrix. * Add the q4-1 kernels. * Proper block sizes. * More details in the tests.	2024-04-15 08:32:47 +02:00
Laurent Mazare	4ecedb1598	Add the full quantized matmul kernels for cuda. (#2057 )	2024-04-14 17:52:08 +02:00
Laurent Mazare	2ac302a5d1	Add the rope THD kernel. (#2014 ) * Add the rope THD kernel. * Cuda kernel for rope-thd. * Add the metal kernels. * Add a dedicated test.	2024-04-05 08:32:58 +02:00
Thomas Santerre	c5626b8271	Add support for "sign" on tensors (#2012 ) * add the sign unary operator * remove uneeded import * remove uneeded import * undo formatting * undo formatting * remove unnecessary redefintion * allow gradient to flow through for sign and round * fix cpu ops to ensure that negzero and positive zero are handled properly * clippy fixes * Properly avoid gradient tracking. * Use a branchless version. --------- Co-authored-by: laurent <laurent.mazare@gmail.com>	2024-04-04 22:32:47 +02:00
Laurent Mazare	f76bb7794a	Bumping the version number to 0.5.0. (#2009 )	2024-04-04 17:48:45 +02:00
Laurent Mazare	318d143224	Relax the contiguous check for cuda kernels. (#2000 ) * Relax the contiguous check for cuda kernels. * Ensure contiguity for RNNs. * Unrelated fix for segment anything. * Better error message + allow concatenating empty slices.	2024-04-03 09:02:38 +02:00
Laurent Mazare	cd29c7ccd4	More ggml cuda kernels (#1977 ) * Add more cuda kernels for quantized matmul. * Add the vec-dot bits. * Expose the quantized matmul-vec kernels. * Also include the quantize-q8-1 kernel. * Glue code for the q8-1 quantization. * mm-vec product via q8-1 quantization. * Add a test. * Add a mm test. * Get the test to return some sensible results. * Also test dmmv. * Fix the launch params. * Allow for tweaking the force_dmmv parameter while it's experimental.	2024-04-01 00:15:48 +02:00
Laurent Mazare	13ae5a34c7	Ensure that the kernels get rebuilt on cuh changes. (#1954 )	2024-03-28 06:56:48 +01:00
Laurent Mazare	196765e995	Use the new rope kernel in mistral. (#1937 ) * Use the new rope kernel in mistral. * Compute the cos and sin with full precision. * Bugfix.	2024-03-25 23:26:05 +01:00
Laurent Mazare	e7f8e72588	Contiguous variant of the rope kernel. (#1929 ) * Contiguous variant of the rope kernel. * Add the cuda kernel. * Metal kernel.	2024-03-25 09:11:20 +01:00
Laurent Mazare	1b98f84a2b	Fast kernels for rotary embeddings. (#1928 ) * Fast kernels for rotary embeddings. * Add a test for the fast CPU kernel. * Rope cuda bindings. * Cuda kernel. * Metal kernel (part 1). * Cuda kernels. * Finish the metal kernel. * Use the new kernels in the quantized example. * Fix warning.	2024-03-24 22:48:52 +01:00
yinqiwen	790037390c	Add cast_bf16_x/cast_x_bf16 when CUDA_ARCH<800 but CUDA_VERSION >= 11000 (#1919 ) - it make possible to load bf16 models on T4(sm75)	2024-03-23 13:44:10 +01:00
Daniël de Kok	fc1fe5e45b	Support scatter/index_add with i64 indices for f16 (#1915 )	2024-03-22 11:51:41 +01:00
Laurent Mazare	af7f8b87d3	Custom op for RmsNorm (#1890 ) * Trying out a custom RmsNorm cuda kernel. * CPU implementation for rms-norm. * Cuda wrappers. * Add some validation. * Add some testing. * More testing.	2024-03-21 06:36:28 +01:00
Laurent Mazare	b219903d0f	Cuda backend optimization (#1886 ) * Attempt at making the kernel faster. * Also adapt the cast kernels. * Also apply to binary ops.	2024-03-20 18:32:55 +01:00
Laurent Mazare	ce9fbc3682	Optimize the cat operation on contiguous tensors (#1855 ) * Add a specialized kernel for copy2d. * Move the cat operations. * Avoid transpositions in cat. * Bugfix. * Bugfix for the cuda kernel. * Add a benchmark. * Add more testing. * Test fix. * Faster kernel. * Add the missing kernel. * Tweak the test. * Add a metal kernel. * Fix for the metal kernel. * Get the tests to pass on metal. * Also use this opportunity to fix the metal kernel for ELU. * Add some bf16 kernels. * Clippy fixes.	2024-03-17 10:49:13 +01:00
Laurent Mazare	e7fc1daa21	Bump the crate versions to 0.4.2. (#1821 )	2024-03-08 22:01:51 +01:00
Laurent Mazare	bd9ab9bc04	Add a cuda kernel for dequantizing q8_0. (#1804 )	2024-03-05 09:50:37 +01:00

1 2 3

134 Commits