Replies: 4 comments 3 replies
|
@OWKenobi I've had simlar thoughts about this. I am curious what others think. From what I know about loras, loras as not so good at /memory/ as they are about learning patterns. So, you're better off feeding a lora /examples/ for it to emulate, rather than data for it to recall. Vector databases like Pinecone seem to be the current state of the art. However from what I gather, they will return relevant hits, and then the raw text of those hits is included in the context. This seems super wasteful to me - Since the AI is "thinking" in vectors, ideally there'd be some way to retrieve hits based on those vectors, and embed them back into the model. @wawawario2 Do you have any thoughts regarding long term memory using lora or embedding, or some way to do it without sacrificing context? |
|
Actively training conversation history into the model, LoRA or otherwise, might be a bit painful, for several reasons. One of which, as described above, is just that LoRAs have to have a very high rank param to even start learning information directly. The other of which is whether high-rank LoRA or genuine finetune... learning information isn't the same as conversational history. It might be a challenge to feed information to a trained model such that it understands to associate that information with the conversational agent. Another being - well LoRA training is a slow and intensive process that you're going to have to fire off every 2048 tokens of conversation, which would get pretty annoying for an active conversation. In terms of solution options: A: Memory toolsets like the one tensiondriven mentioned above are a good tool, if designed well to activate automatically and relevantly. B: Parallel Context Windows is a very interesting looking potential project being researched that may be able to achieve more input room - #540 C: This was implemented for Stable Diffusion: https://arxiv.org/abs/2302.05543 and I think the core ideas may be applicable here as well. Essentially the idea is you can train a small module on top of an existing model, like a hypernetwork, but specifically designed to take an additional input buffer and bias the output based on the additional input. I think it should be possible to build a ControlNet-style module over top of LLaMA that is trained specifically to take an efficient compression of a large additional context window and bias the base model to make use of it, and it should in theory be possible to train it more efficiently than training the model itself on large context windows. |
|
I had this similar idea and I'm researching about it these days. If you are doing similar research I would be happy to help. |
|
Is there a place where I can download lora or is it an ongoing project. While it's easy to find lora for stable diffusion, I think it's too early for text ai. I really want to see lora for Text Ai |
Uh oh!
There was an error while loading. Please reload this page.
Just wanted to share and refine ideas with your comments:
My long-term idea is:
The bonus of this approach: when, in the future, a stronger, better model gets released, you can take your "personality" with you and continue the chat with a different bot.
On a business side of view, you can fine tune your model just by talking to it, and when you feel it is good enough, simply let it sleep, and it will put all info in the long-term-memory-lora. You can repeat this process unlimited times without compromising on quality or speed.
Todo
has anyone ever tried to use a personal chat to train a lora? Does it work good enough? How long does it take? (let's assume money for hardware is not an issue)
All reactions