candle

mirror of https://github.com/huggingface/candle.git synced 2025-06-16 10:38:54 +00:00

Author	SHA1	Message	Date
Ivar Flakstad	7ec4f64d38	Attempt at fixing M1/M2 metal async copy bug	2024-09-06 15:59:35 +02:00
Ivar Flakstad	8712ceb84f	Revert "slight changes to async ops" This reverts commit `16b49dfd89`.	2024-09-02 12:34:11 +02:00
Ivar Flakstad	f9b2bb4d46	Revert "Alter metal simdgroup matrix load/store ops" This reverts commit `560f666d29`.	2024-09-02 12:34:09 +02:00
Ivar Flakstad	aefca7f8e6	Revert "Testing ushort intermediate in case combo of async and f16/bf16 is the issue" This reverts commit `c3b0757995`.	2024-09-02 12:33:58 +02:00
Ivar Flakstad	c3b0757995	Testing ushort intermediate in case combo of async and f16/bf16 is the issue	2024-08-24 16:59:33 +02:00
Ivar Flakstad	560f666d29	Alter metal simdgroup matrix load/store ops	2024-08-21 08:57:50 +02:00
Ivar Flakstad	16b49dfd89	slight changes to async ops	2024-08-08 12:30:22 +02:00
Ivar Flakstad	fc0deede31	See if M1/M2 likes metal sources better than metallibs compiled on M3	2024-08-08 09:32:13 +02:00
Ivar Flakstad	e178aacead	Re-revert the reverted revision (bf16 gemm metal)	2024-08-06 12:37:36 +02:00
Ivar Flakstad	bb191c25d5	Update metallib with device matrix offsets	2024-08-05 15:05:00 +02:00
Laurent Mazare	2be9bd211e	Support for mistral-nemo. (#2396 )	2024-08-04 19:52:40 +02:00
Laurent Mazare	89eae41efd	Support the flux-dev model too. (#2395 )	2024-08-04 12:16:24 +02:00
MilkFather	c0a559d427	optimize gradient for silu a bit (#2393 )	2024-08-04 11:24:17 +02:00
Laurent Mazare	aa7ac1832d	Simplify handling of flux modulations. (#2394 )	2024-08-04 11:09:54 +02:00
Laurent Mazare	19db6b9723	Add the flux model for image generation. (#2390 ) * Add the flux autoencoder. * Add the encoder down-blocks. * Upsampling in the decoder. * Sketch the flow matching model. * More flux model. * Add some of the positional embeddings. * Add the rope embeddings. * Add the sampling functions. * Add the flux example. * Fix the T5 bits. * Proper T5 tokenizer. * Clip encoder path fix. * Get the clip embeddings. * No configurable weights in layer norm. * More weights related fixes. * Yet another shape fix. * DType fix. * Fix a couple more shape issues. * DType fixes. * Fix the latent dims. * Fix more shape issues. * Autoencoder fixes. * Get some generations out. * Bugfix. * T5 padding. * Clippy fix. * Add the decode only mode. * Fix. * More fixes. * Finally get some generations to work. * Add readme.	2024-08-04 08:14:33 +02:00
Laurent Mazare	0fcb40b229	Revert the bf16 gemm metal changes for now. (#2386 )	2024-08-01 23:08:47 +02:00
Justin Sing	6991a37b94	update: LSTMState and GRUState fields to be public (#2384 )	2024-08-01 16:30:32 +02:00
Laurent Mazare	9ca277a9d7	Fix cargo fmt. (#2383 ) * Fix cargo fmt. * Clippy fix. * Cosmetic tweaks.	2024-08-01 14:19:41 +02:00
Joan Fontanals	2e9c010609	Jina Bert Example fix and more configuration (#2191 ) * fix: fix jina bert example logic * feat: enable jina embeddings de * feat: allow more flexibility on Jina Bert	2024-08-01 13:59:20 +02:00
Jani Monoses	ac51f477eb	Add Hiera vision model. (#2382 )	2024-08-01 11:59:22 +02:00
Laurent Mazare	d4b6f6eef6	Add a minimal test for the metal bf16 matmul. (#2381 )	2024-08-01 11:22:46 +02:00
Laurent Mazare	957d604a78	Enable BF16 on metal. (#2380 )	2024-08-01 11:05:07 +02:00
Takanori MAEHARA	ce90287f45	Add get_ids to GradStore (#2379 )	2024-08-01 10:56:13 +02:00
Laurent Mazare	1ba87a9450	Use BF16 on metal when possible. (#2378 )	2024-08-01 10:48:58 +02:00
Yun-Jhong Wu	bd80078acf	Fix log_sum_exp to handle large positive/negative inputs (#2367 )	2024-08-01 10:37:02 +02:00
ivarflakstad	fea46cb719	Metal bgemm min changes (#2364 ) * Add updated mfa metallib * Add bgemm and tests	2024-08-01 10:06:04 +02:00
Laurent Mazare	8696cf6494	Enable the affine kernel for u8/u32. (#2376 )	2024-08-01 10:03:11 +02:00
Zheng Li	4a52aeb437	bert attention mask (#1934 ) * bert attention mask * Allow for using None as a mask. * Revert part of the changes so that the proper default mask applies. * Cosmetic change. * Another cosmetic tweak. --------- Co-authored-by: Laurent <laurent.mazare@gmail.com>	2024-08-01 08:26:19 +02:00
ivarflakstad	24d54d0ff9	Bump image crate version so ImageReader is available without aliasing (#2365 )	2024-07-29 17:41:33 +02:00
Jacob Marshall	636eff652a	change DTypes (fixes #2355 ) (#2363 )	2024-07-28 14:36:05 +02:00
Eric Buehler	0f5cbb08b3	Add support for Llama 3.1 (#2359 ) * Add Llama 3.1 rope * Clippy * Format * Clippy * Add support for multiple eos tokens: * Untagged either * Remove either dep and fix settings.json * Make the max positional embeddings configurable	2024-07-26 21:32:26 +02:00
Laurent Mazare	ddafc61055	Use RAII for terminating the encoding. (#2353 )	2024-07-24 16:29:56 +02:00
Laurent Mazare	a925ae6bc6	Use a trait for the encoder provider (so that encoder can ultimately be reused). (#2352 )	2024-07-24 09:27:30 +02:00
shua	6056fd5c90	onnx: fix pad, unsqueeze (#2317 ) * onnx: fix pad, unsqueeze both implementations have off-by-one errors: - Pad 'reflect' cycle for eg `dim==3` is `[0,1,2,1]` which has length of 4 (or `dim2 - 2`) not 5 (current code `dim2 - 1`) - Unsqueeze(-1) for tensor with `dim==3` should be 3 (ie `dim+index+1`) not 2 (ie currently `dim+index`) in addition, Pad is incorrectly calculating the starting padding. If we want to pad out 2 elements to the start, and we have this cycle of indices of length 6, then we should skip 4 elements, but currently we skip 2. A more visual representation of what's going on is below: ``` pad_start: 2 data: [a,b,c,d] indices: [0, 1, 2, 3, 2, 1, 0, 1, 2, 3, 2, 1, 0, ..] // zigzag between 0..4 actual: skip [ c d\| c b a b] expected: ~ skip ~ [ c b\| a b c d] ``` The values between `[` and `\|` are padding and the values between `\|` and `]` in the example should match the original data being padded. * Fix clippy lints. --------- Co-authored-by: Laurent <laurent.mazare@gmail.com>	2024-07-23 23:10:57 +02:00
Caio Petrucci Rosa	ebc9aa60bc	fix clip example title (#2345 )	2024-07-23 22:55:18 +02:00
donjuanplatinum	2489a606fe	feat(candle-transformers/models/codegeex4-9b): add codegeex4-9 (#2334 ) * feat(candle-transformers/models/codegeex4-9b): add codegeex4-9b transoformers * change mod.rs * feat(candle-examples/codegeex4-9b) * Update codegeex4_9b.rs * Update main.rs * Update codegeex4_9b.rs * Update main.rs * fmt * fix * fmt * Clippy fix. * Remove some print statements. * Avoid using unwrap. * 1. add README 2. change the print fmt * Another clippy fix. --------- Co-authored-by: Laurent <laurent.mazare@gmail.com>	2024-07-21 13:00:41 +02:00
Laurent Mazare	3c815b1dca	Pin the revision used by moondream. (#2340 )	2024-07-18 10:49:46 +02:00
Laurent Mazare	42891cc613	Add mathstral in the examples. (#2339 )	2024-07-18 08:24:49 +02:00
Ivor Wanders	f25173d68b	Fix for backprop in ConvTranspose2D with stride of 2 (#2337 ) * Add gradient test for conv_transpose2d with stride of 2. * Swap dilation and stride in ConvTranspose2D backpropagation. Without this, a shape mismatch occurs with a stride of 2 and dilation of 1. * Add further tests of the ConvTranspose2D gradient. Values calculated with torch, minor numerical errors adjusted and commented.	2024-07-17 19:22:23 +02:00
Alexey Gerasev	6a4741bbf9	Fix Elu gradient NaN on large input (#2328 ) * Fix Elu gradient NaN on large input * Reuse previously computed exp in Elu	2024-07-16 14:41:16 +02:00
Laurent Mazare	30cdd769f9	Update the flash attn kernels. (#2333 )	2024-07-15 20:37:36 +02:00
Josh Collyer	d74fbed334	Pinning cudarc to 0.11.6 (#2332 )	2024-07-15 15:29:08 +02:00
Zhuo Jinggang	c63048d374	add quantized qwen2 (#2329 ) * add quantized version of qwen2 and corresponding example for qwen2-instruct * fix quantized qwen2 clippy error	2024-07-12 10:00:03 +02:00
Jani Monoses	a226a9736b	Add Mobilenet v4 (#2325 ) * Support different resolutions in load_image() * Added MobilenetV4 model. * Add MobileNetv4 to README	2024-07-09 13:52:20 +02:00
Laurent Mazare	25960676ca	Add a basic metal example with capture (#2324 ) * Add some tracing. * Get the trace to work.	2024-07-09 12:38:11 +02:00
v-espitalier	9cd54aa5d4	Add EVA-02 model ( https://arxiv.org/abs/2303.11331 ) (#2311 ) * Add EVA-02 model ( https://arxiv.org/abs/2303.11331 ) * Clippy fix. * And apply fmt. --------- Co-authored-by: v-espitalier <> Co-authored-by: Laurent <laurent.mazare@gmail.com>	2024-07-07 20:09:31 +02:00
shua	eec11ce2ce	onnx: implement Size op (#2316 )	2024-07-07 19:56:36 +02:00
shua	9182f9f5c2	ignore editor config folders (#2315 )	2024-07-07 19:43:48 +02:00
v-espitalier	ecff05d72b	Beit: Add the gen_relative_position_index() function (#2306 ) Co-authored-by: v-espitalier <>	2024-07-04 09:45:26 +02:00
v-espitalier	7f1ba8038c	Add Beit model ( https://arxiv.org/abs/2106.08254 ) (#2305 ) Co-authored-by: v-espitalier <>	2024-07-01 22:11:48 +02:00

1 2 3 4 5 ...

2124 Commits