* make attention faster for a couple models * remove unused generation flags * add comment on lora * include text files as well |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| base.py | ||
| cohere.py | ||
| gemma.py | ||
| layers.py | ||
| llama.py | ||
| mixtral.py | ||
| olmo.py | ||
| phi.py | ||
| phixtral.py | ||
| plamo.py | ||
| qwen.py | ||
| qwen2.py | ||
| stablelm.py | ||
| starcoder2.py | ||